
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
NVIDIA released performance data showing the Vera Rubin NVL72 system's efficiency in handling token-intensive agentic AI workloads.
Why it matters
This allows power-constrained AI factories to perform more agentic work for the same energy footprint while reducing the cost per million tokens.
The details
- Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than GB300 NVL72.
- The system enables up to 35x lower token cost than GB300 NVL72.
- Agentic AI workloads consume 15x more tokens than simple chat requests.
Show entities and relationshipsHide entities and relationships
In this article
Products
Technologies
Topics
Companies
Key connections
NVIDIA owns GB300 NVL72
NVIDIA designs and manufactures the GB300 NVL72 rack system.
NVIDIA owns NVIDIA Hopper
NVIDIA developed the Hopper GPU computing architecture.
NVIDIA owns Spectrum-6 SPX
NVIDIA builds the Spectrum-6 SPX networking switch for AI factories.
NVIDIA owns TensorRT LLM
NVIDIA develops the TensorRT LLM inference runtime.
NVIDIA develops the MegaMoE fused CUDA kernel.
SemiAnalysis owns AgentX
SemiAnalysis created and maintains the AgentX benchmark workload for agentic AI.
Show 34 more connectionsShow fewer connections
Groq owns Groq 3 LPU
Groq designs the Groq 3 LPU processor integrated into next-generation AI platforms.
NVIDIA Vera Rubin NVL72 competes with GB300 NVL72
Vera Rubin NVL72 succeeds GB300 NVL72, offering up to 30x higher throughput per megawatt.
GB300 NVL72 competes with NVIDIA Hopper
GB300 NVL72 delivers up to 15x better throughput per megawatt than the Hopper architecture.
NVIDIA Vera Rubin NVL72 is built with NVIDIA Vera CPU
Vera Rubin NVL72 incorporates the Vera CPU in its seven-chip architecture.
NVIDIA Vera Rubin NVL72 is built with NVIDIA Rubin GPU
Vera Rubin NVL72 utilizes Rubin GPUs for accelerated compute and inference.
NVIDIA Vera Rubin NVL72 is built with Groq 3 LPU
Vera Rubin NVL72 integrates the Groq 3 LPU into its seven-chip platform.
NVIDIA Vera Rubin NVL72 is built with NVIDIA NVLink 6 Switch
Vera Rubin NVL72 uses NVLink 6 Switches for its scale-up domain interconnect.
NVIDIA Vera Rubin NVL72 is built with NVIDIA BlueField-4 DPU
Vera Rubin NVL72 includes the BlueField-4 DPU for infrastructure processing.
NVIDIA Vera Rubin NVL72 is built with Spectrum-6 SPX
Vera Rubin NVL72 connects via Spectrum-6 SPX networking switches.
NVIDIA Vera Rubin NVL72 is built with NVIDIA ConnectX-9 SuperNIC
Vera Rubin NVL72 integrates ConnectX-9 SuperNICs for network scaling.
NVIDIA Vera Rubin NVL72 uses TensorRT LLM
Vera Rubin NVL72 leverages TensorRT LLM runtime for inference acceleration.
NVIDIA Vera Rubin NVL72 uses NVIDIA Dynamo
Vera Rubin NVL72 runs on NVIDIA Dynamo serving framework for hardware-software codesign.
NVIDIA Vera Rubin NVL72 uses MegaMoE
Vera Rubin NVL72 executes MegaMoE fused CUDA kernels for efficient mixture-of-experts computation.
GB300 NVL72 is built with NVIDIA Blackwell
GB300 NVL72 is built on the NVIDIA Blackwell platform architecture.
AgentX uses DeepSeek V4 Pro
AgentX benchmark evaluates inference throughput using the DeepSeek V4 Pro model.
AgentX benchmark tests agentic performance across models including Kimi K3.
AgentX uses MiniMax M3
AgentX benchmark tests agentic performance across models including MiniMax M3.
AgentX benchmark tests agentic performance across models including GLM5.3.
AgentX benchmark tests agentic performance across models including Qwen3.5.
DeepSeek V4 Pro competes with Kimi K3
DeepSeek V4 Pro and Kimi K3 are evaluated as competing agentic AI models.
DeepSeek V4 Pro competes with MiniMax M3
DeepSeek V4 Pro and MiniMax M3 are competing agentic models in benchmark evaluations.
DeepSeek V4 Pro competes with GLM5.3
DeepSeek V4 Pro and GLM5.3 are competing agentic AI models.
DeepSeek V4 Pro competes with Qwen3.5
DeepSeek V4 Pro and Qwen3.5 are competing agentic models.
OpenRouter uses Agentic AI
OpenRouter collects and analyzes workload data on agentic AI token consumption.
NVIDIA Vera Rubin NVL72 uses Disaggregated Serving
Vera Rubin NVL72 enables disaggregated serving to separate prefill from decode stages.
NVIDIA Vera Rubin NVL72 uses Rate Matching
Vera Rubin NVL72 applies rate matching to synchronize prefill and decode token generation.
NVIDIA Vera Rubin NVL72 uses Large-Scale Expert Parallelism
Vera Rubin NVL72 uses large-scale expert parallelism across its GPU domain for MoE models.
NVIDIA Vera Rubin NVL72 uses Distributed KV-Caching
Vera Rubin NVL72 extends memory across the GPU domain using distributed KV-caching.
NVIDIA Vera Rubin NVL72 uses KV-Cache Offloading
Vera Rubin NVL72 tiers less-active context to host and storage with KV-cache offloading.
NVIDIA Vera Rubin NVL72 uses KV-Aware Routing
Vera Rubin NVL72 routes incoming requests to GPUs holding relevant cached context.
NVIDIA Rubin GPU uses NVFP4 Quantization
Rubin GPUs utilize NVFP4 quantization to compress model weights to 4-bit precision.
NVIDIA Vera Rubin NVL72 is built with NVIDIA NVLink
Vera Rubin NVL72 is interconnected using NVLink technology for high-bandwidth communication.
NVIDIA is related to DSX MaxLPS
NVIDIA develops DSX MaxLPS technologies to manage power across GPU and rack levels.
NVIDIA is a partner of SemiAnalysis
NVIDIA collaborated with SemiAnalysis to benchmark inference throughput on the AgentX workload.
Related events
NVIDIA Demonstrates Vera Rubin NVL72 Agentic AI Performance Gains
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.