A claim about systems, not simply chips
NVIDIA’s assertion that Rubin can handle up to ten times more AI agents than Blackwell is broadly grounded in the company’s latest description of its Vera Rubin platform. Yet the claim needs careful interpretation. It does not mean that any single Rubin GPU will automatically run ten times as many agents as every Blackwell GPU in every deployment.
The stated comparison is between the Vera Rubin platform and the preceding Grace Blackwell platform at scale. NVIDIA frames the result as “agentic throughput” per unit of energy, using an internal two-trillion-parameter mixture-of-experts workload. In other words, the measurement concerns how much useful work a large, multi-step AI system can sustain within a power budget. It is a performance-and-efficiency claim about an integrated data-centre configuration, not a simple benchmark for one model query or a consumer graphics card.
That distinction matters because an AI agent is not equivalent to a conventional chatbot exchange. A typical agent may plan a task, retrieve data, invoke software tools, write or execute code, check intermediate results and repeat the process. Its computing requirements are therefore shaped by response speed, context length, memory capacity, network latency and the coordination of many concurrent processes, as well as by raw accelerator performance.
Why agentic workloads expose more bottlenecks
NVIDIA’s design rationale is that large agentic workloads are limited by much more than arithmetic throughput. Models with long contexts and mixture-of-experts architectures repeatedly move data between high-bandwidth memory, compute units and multiple processors. Small delays can compound when one agent waits for a tool call or when a group of agents shares a distributed model.
Rubin is intended to address these constraints through hardware and system co-design. The Rubin GPU introduces HBM4 memory, a third-generation Transformer Engine and changes aimed at reducing latency between dependent computing kernels. NVIDIA also highlights compression and sparsity techniques intended to reduce unnecessary movement of data and calculation.
The wider Vera Rubin configuration is equally important to the headline number. Its NVL72 rack combines 72 Rubin GPUs with 36 Vera CPUs, high-speed NVLink switching, networking components and liquid cooling. The Vera CPU capacity is relevant because agents need conventional compute for orchestration, evaluation, tool execution and sandboxed tasks; not all of that work belongs on a GPU.
This is why the comparison is more accurately read as a claim about an AI factory: a coordinated cluster of processors, memory, networking, cooling and power-management equipment that is tuned for sustained inference.
Efficiency is central to the promised gain
The most notable aspect of NVIDIA’s figure is that it is expressed per unit of energy. Data-centre operators increasingly face constraints in electrical supply, cooling capacity and the cost of bringing new facilities online. Under those conditions, a platform that produces more completed agent steps or tokens for the same megawatt can be commercially more consequential than a narrow peak-compute statistic.
NVIDIA says Vera Rubin uses power smoothing to reduce sharp demand spikes, enabling infrastructure to make better use of its installed electrical capacity. It also says its DSX MaxLPS approach can permit up to 40% more GPUs to be provisioned in a fixed power envelope under suitable operating conditions. Such system-level methods help explain why the claimed improvement cannot be attributed solely to the Rubin GPU architecture.
The company separately says Rubin can reduce inference cost per token by up to tenfold compared with Blackwell, and train some mixture-of-experts models with four times fewer GPUs. These are related but distinct claims. Lower token cost, more agent throughput per watt and fewer GPUs for training each depend on the chosen model, precision, software stack and deployment design.
What remains to be proven independently
NVIDIA’s performance figures are useful indicators of its product positioning, but they should not be treated as independently validated industry-wide results. The tenfold agent-throughput metric is based on NVIDIA’s internal workload, and its real-world outcome will vary with model architecture, context length, agent framework, tool-use patterns and utilisation levels.
For example, an enterprise running relatively short, single-agent workflows may see different benefits from an organisation operating a massive reasoning model with thousands of concurrent agents. A deployment limited by storage, external APIs, databases or poorly parallelised software may not fully benefit from accelerator and interconnect improvements. Conversely, applications that spend substantial time generating tokens, maintaining long contexts and coordinating many sub-agents could be closer to the workloads Rubin was designed to improve.
Independent benchmarking will therefore be important once systems are broadly available. Tests should disclose the model, precision, prompt and context lengths, concurrency, latency targets, total system power, cooling assumptions and software versions. They should also compare cost and reliability rather than focusing only on the maximum number of agents launched at once.
A shift in the unit of competition
Rubin illustrates NVIDIA’s broader shift from selling a GPU story to selling a rack- and pod-scale infrastructure story. The company is presenting the data centre itself as the unit of compute, with CPUs, GPUs, switching, storage, security and power delivery designed as a single platform.
That approach is well aligned with the rise of agentic AI, where useful performance depends on keeping a long chain of reasoning, retrieval and tool execution moving without bottlenecks. The “ten times more agents” headline captures the ambition, but its proper meaning is narrower: NVIDIA is claiming up to a tenfold improvement in energy-normalised agentic throughput for a particular large-scale Vera Rubin configuration and workload relative to Grace Blackwell. The practical test will be whether customers can reproduce meaningful portions of that gain in production systems.
Sources
- Nvidia Rubin zvládne až 10× více AI agentů než Blackwell — Svět hardware
- Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI — NVIDIA Technical Blog
- NVIDIA Vera Rubin Ramps Into Full Production to Power Agentic AI Factories Worldwide — NVIDIA Newsroom
- Infrastructure for Scalable AI Reasoning | NVIDIA Vera Rubin Platform — NVIDIA



