Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
The landscape of artificial intelligence is shifting rapidly from simple, transactional chatbots to complex, autonomous agents. According to recent data from OpenRouter, this transition comes with a significant computational price tag: agentic AI workloads consume roughly 15 times more tokens than standard chat requests.
This surge in demand is a natural byproduct of how agents operate. When an AI agent is tasked with a complex objective—such as conducting a deep-dive financial analysis on a potential investment—it does not simply provide an answer. It queries financial databases, parses news filings, triggers sub-agents to perform peer comparisons, and synthesizes multi-layered valuations. Because each step in this reasoning chain relies on the output of the previous one, the context window grows exponentially. As this agentic paradigm becomes the industry standard for everything from software engineering to customer support, the underlying infrastructure must evolve to handle this massive, sustained token demand.
Redefining Efficiency: The Vera Rubin Advantage
NVIDIA has unveiled new performance benchmarks for its Vera Rubin NVL72 systems, revealing a massive leap in efficiency. When measured against the NVIDIA GB300 NVL72, the Vera Rubin architecture delivers up to 30x higher throughput per megawatt for agentic workloads.
These figures were derived using the SemiAnalysis AgentX workload, a benchmark that captures real-world agentic coding sessions. By preserving the nuances of actual context growth, tool calls, and sub-agent spawning, the test provides a realistic look at how these systems perform in production environments. For data centers and AI factories operating under strict power constraints, this advancement is transformative, effectively enabling 30 times more agentic output for the same energy footprint.
"For power-constrained AI factories, throughput per megawatt determines AI factory revenue and cost per million tokens determines the profit margin on that revenue."
Moving Beyond Simple Inference
Traditional performance metrics, which often focus on single-request latency for short documents, are no longer sufficient for the agentic era. In these workflows, context accumulates across numerous steps, often reaching hundreds of thousands of input tokens.
NVIDIA’s testing demonstrates that the Blackwell platform provides a robust foundation for this new reality. Across a variety of advanced models—including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5, and DeepSeek V4 Pro—the platform shows industry-leading efficiency. Notably, the GB300 NVL72 already offers a 15x improvement in throughput per megawatt over the previous Hopper architecture. The Vera Rubin platform pushes this boundary even further, elevating the entire performance curve to achieve its 30x efficiency gain on the DeepSeek V4 Pro model.
Beyond raw throughput, the economic impact is clear: Vera Rubin NVL72 can reduce the cost per million tokens by up to 35x compared to the GB300 NVL72. This drastic reduction in cost is essential for businesses looking to run autonomous agents continuously at scale.
Extreme Codesign: The Architecture of Scale
The performance gains seen in the Vera Rubin NVL72 are not the result of a single hardware tweak, but rather a philosophy of "extreme codesign" that spans the entire platform stack. NVIDIA has integrated several critical optimizations to ensure that hardware and software work in perfect harmony:
- Disaggregated Serving: Separates context processing (prefill) from response generation (decode), allowing each to scale independently based on demand.
- Large-Scale Expert Parallelism: Distributes expert sub-networks across the entire GPU domain, optimizing the performance of mixture-of-experts models.
- Distributed KV-Caching: Extends memory across the scale-up domain while offloading less-active context to host storage, preventing redundant recomputations.
- KV-Aware Routing: Intelligently directs incoming requests to GPUs that already possess the relevant cached context.
- Fused CUDA Kernels: Technologies like MegaMoE combine computation and inter-GPU communication into single execution passes, minimizing idle time.
- NVFP4 Quantization: Compresses model weights to 4-bit precision, significantly reducing memory footprints while maintaining output quality.
These hardware-level advancements are supported by the sixth-generation NVLink interconnect and NVLink Switches, which provide 10x higher packet rates and 3x lower latency compared to standard Ethernet solutions.
A Holistic AI Factory Ecosystem
While the Vera Rubin NVL72 is the star of this announcement, it is part of a broader, seven-chip architecture designed for the future of AI factories. The full platform includes the NVIDIA Vera CPU, Groq 3 LPU, BlueField-4 DPU, Spectrum-6 SPX, and ConnectX-9 SuperNIC.
By integrating these components with the NVIDIA TensorRT LLM inference runtime and the Dynamo serving framework, NVIDIA is providing a comprehensive stack that is already in production and scaling across the global ecosystem. As these systems continue to benefit from ongoing software optimizations, the gap between traditional inference and agentic-scale performance is set to widen, cementing the Vera Rubin NVL72 as the new benchmark for the next generation of AI-driven productivity.