Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
NVIDIA's new Vera Rubin NVL72 systems demonstrate up to 30x higher throughput per megawatt and 35x lower token cost compared to GB300 NVL72 for agentic AI workloads. This significant efficiency leap addresses the high token consumption characteristic of complex AI agent t…
Intelligence analysis by Gemini 2.5 Flash

AI agents, which perform multi-step tasks like financial research or software development, consume vastly more tokens than simple chat requests due to accumulating context. NVIDIA's Vera Rubin NVL72 platform is engineered to meet this demand with unprecedented efficiency, offering a substantial performance boost per unit of energy for power-constrained AI factories.
Imagine you have a super-smart robot helper that needs to do many steps to finish a big job, like researching a whole company. Each step makes the job bigger, like adding more pages to a book. NVIDIA has made a new super-computer brain, called Vera Rubin NVL72, that helps these robot helpers work much, much faster and use way less electricity. It's like giving your robot helper a super-efficient brain that can read and think 30 times quicker while barely using any battery power, making big jobs much easier and cheaper to do.
Analysis
Vera Rubin NVL72
NVIDIA's Vera Rubin NVL72 systems represent a significant leap in efficiency for AI agent workloads, demonstrating up to 30 times higher throughput per megawatt than the NVIDIA GB300 NVL72. This performance gain is crucial for AI factories facing power constraints, directly translating into more agentic work for the same energy footprint. Furthermore, the Vera Rubin NVL72 achieves up to 35 times lower cost per million tokens compared to its predecessor, making continuous, large-scale agent operations more economically viable.
These early results, measured using the SemiAnalysis AgentX workload, highlight NVIDIA's accelerated pace of innovation in AI infrastructure. The platform's enhanced capabilities are designed to support the demanding, long-context requirements of agentic AI, ensuring that performance continues to improve with ongoing software optimizations across both Vera Rubin and GB300 NVL72 systems.
Agentic Workloads
Agentic AI workloads differ fundamentally from simpler tasks like chat or document summarization, which typically involve input and output sequences ranging from 1K to 8K tokens. In contrast, agentic sessions accumulate context across multiple steps, often reaching hundreds of thousands of input tokens with wide variability in both input and output lengths. This necessitates a new approach to performance measurement that captures the entire agent workflow rather than just a single inference request.
NVIDIA measured this inference throughput data using the SemiAnalysis AgentX workload, which consists of recorded real-world agentic coding sessions. This benchmark accurately preserves actual context growth, tool calls, and sub-agent spawning, providing a realistic assessment of performance. The NVIDIA Blackwell platform, including GB300 NVL72, already delivers leading performance across various agentic models such as Kimi K3, MiniMax M3, GLM5.3, Qwen3.5, and DeepSeek V4 Pro, with Vera Rubin further extending this advantage.
Extreme Codesign
NVIDIA Vera Rubin NVL72 achieves its multifold performance gains through extreme codesign across every layer of the platform, integrating modern inference optimization techniques. Disaggregated serving, for instance, separates context processing (prefill) from response generation (decode) to allow independent scaling and rate matching synchronizes their speeds for maximum efficiency. Large-scale expert parallelism distributes sub-networks of mixture-of-experts models across the GPU domain, while distributed KV-caching extends memory and offloads less-active context to host and storage, preventing recomputation.
Key hardware and software innovations include NVIDIA Rubin GPUs' enhanced fifth-generation Tensor Cores and third-generation Transformer Engine, which accelerate both prefill and decode stages. NVFP4 quantization compresses model weights to 4-bit precision, reducing memory footprint and boosting throughput without sacrificing output quality. The NVL72 scale-up domain, powered by sixth-generation NVLink interconnect technology and NVLink Switches, provides the high-bandwidth and low-latency inter-GPU communication essential for these advanced techniques, all supported by an optimized software stack including NVIDIA TensorRT LLM and NVIDIA Dynamo.
Key points
- NVIDIA Vera Rubin NVL72 systems offer up to 30x higher throughput per megawatt for agentic AI workloads.
- The new platform delivers up to 35x lower cost per million tokens compared to GB300 NVL72.
- Agentic AI workloads consume significantly more tokens due to accumulating context across multi-step tasks.
- Performance was measured using the SemiAnalysis AgentX workload, reflecting real-world agentic coding sessions.
- Key innovations include disaggregated serving, distributed KV-caching, NVFP4 quantization, and enhanced NVLink interconnect technology.
This significant leap in efficiency for AI agents could accelerate their deployment across various industries, making complex AI tasks more accessible and affordable. It promises to unlock new capabilities for businesses and researchers, enabling more sophisticated automation and deeper insights with reduced operational costs and energy consumption.



