한국어판: 한국어로 읽기
NVIDIA and CoreWeave are pitching Vera Rubin not as a faster chip in isolation, but as infrastructure for a continuous agent loop: run, observe, evaluate, improve and redeploy.
On September 30, 2026, NVIDIA and CoreWeave announced that NVIDIA Vera Rubin NVL72 systems are entering limited production availability on CoreWeave Cloud. Cognition, the company behind Devin, is named as the first production customer.
The launch is notable because the competitive claim is broader than accelerator speed. NVIDIA and CoreWeave are describing an integrated stack of GPUs, networking, CPUs and cloud software designed to support long-running agentic workloads from inference through post-training and evaluation.
The product being sold is increasingly the loop
Traditional AI infrastructure comparisons often focus on training throughput or inference speed. Agentic systems add more moving parts: tool calls, sandboxes, memory, orchestration, evaluation and repeated improvement cycles.
CoreWeave’s message is that Vera Rubin NVL72, Spectrum-X 102.4T Ethernet and the forthcoming Vera CPU can be operated as one production platform. NVIDIA similarly frames the system as a way to close the loop between training, deployment, observation and post-training.
That changes the unit of competition. The question becomes less “which accelerator is fastest?” and more “which infrastructure can keep a large agent system productive across the entire operational cycle?”
The performance numbers need careful attribution
CoreWeave says Cognition is seeing a 4.8× increase in total token throughput on Vera Rubin NVL72, and reports a 3.8× output-token-throughput gain per GPU over GB200 NVL72 at matched interactivity for a reinforcement-learning workload.
Those are vendor-reported figures from CoreWeave’s own deployment and benchmark conditions. They are useful evidence that the system is working in production, but they are not independent proof that every workload or cloud configuration will see the same economics.
Similarly, “production ready” should be read in context. CoreWeave describes Vera Rubin as available to select customers rather than universally available across its entire customer base.
Why the CPU matters in an agent stack
CoreWeave also says it plans to offer NVIDIA Vera, a CPU designed specifically for agent workloads. The argument is that GPUs handle generation and reasoning while CPUs increasingly execute the surrounding work: sandboxes, tool calls and data pipelines.
If that model holds, agent infrastructure becomes more heterogeneous. The workload is split across specialized components rather than forcing every step through the same processor.
The bigger shift
The durable signal is not that one cloud provider has a new rack. It is that AI infrastructure vendors are increasingly optimizing for continuous autonomous work rather than isolated model calls.
As agents become longer-running and more tool-intensive, infrastructure efficiency will depend on the whole feedback loop: inference latency, orchestration overhead, networking, CPU work, post-training and evaluation. That is a different competitive problem from simply counting GPUs.
What remains uncertain is the most important commercial question: whether the performance gains translate into better unit economics across a wide range of real workloads, not just the launch examples.