Huang repositioned Rubin as the foundation of the "agentic AI factory"
After the GTC 2026 keynote, the Vera Rubin platform is in full production, with deliveries starting in the second half of 2026. The most striking official numbers: versus the prior Blackwell generation, inference cost per million tokens drops to one tenth, and tokens per megawatt rise tenfold. This time NVIDIA is not selling a single chip. It is selling a rack-scale system designed end to end for agents.
One rack is one supercomputer
Vera Rubin NVL72 is not a new graphics card. It is 72 Rubin GPUs plus 36 Vera CPUs linked by sixth-generation NVLink into a single "giant GPU." NVIDIA says it trains mixture-of-experts models with one fourth the GPUs of Blackwell and at one tenth the token cost for inference.
Paired with the Groq 3 LPX inference rack, it reaches up to 35 times the throughput per megawatt for trillion-parameter models. The Vera CPU rack alone sustains 22,500 concurrent reinforcement-learning or agent-sandbox environments. CMX context storage offloads the KV cache from GPU memory into a dedicated storage layer, adding up to 5 times tokens per second and 5 times power efficiency. The whole platform is seven chips and five rack types operating together as one AI supercomputer.
Why "per megawatt" now
In 2026, inference has overtaken training as the dominant category of AI compute consumption. A single agent conversation can eat 15 times the tokens of a traditional application. Huang's logic is straightforward: pack more intelligence into the same power footprint, and always-on agents become viable at consumer price points.
NVIDIA aimed the platform at what it calls the four scaling laws: pretraining, post-training, test-time scaling, and agentic scaling. In other words, it is betting that future compute spending is dominated by inference and agents, not training. More than 80 MGX ecosystem partners are building this generation of racks, and CoreWeave, Google Cloud, Microsoft Azure, and Oracle have already deployed or are running them. The rack-scale supply chain spans 30 countries and over 350 factories.
But read the vendor numbers as vendor numbers
The "token cost drops to one tenth" figure is measured on a specific model, Kimi-K2-Thinking, under fixed input and output lengths. Do not treat it as a universal number across all workloads. The 10x and 35x figures are also "projected" and shift with model, batch size, and network topology. The training-side "one fourth the GPUs" likewise depends on the MoE architecture and a specific configuration.
The more real constraint is power and supply chain. No matter how pretty the per-megawatt number, deployment still depends on datacenter power and delivery pace. Full production is real, but "the whole industry gets 10x cheaper tomorrow" is not. Sovereign AI clouds such as Saudi HUMAIN, UAE G42, and India's E2E are already queued, which shows this generation's capacity is a scarce resource, not an open tap.
One overlooked inflection point
When inference cost falls by an order of magnitude, the effect is not only lower prices. Real-time, long-context, and multi-agent scenarios that were uneconomic before suddenly become feasible. In other words, a base like Rubin matters less for being faster and more for making previously unprofitable work ordinary. That is the part of a generational switch worth watching most closely.
What this means for you
The trend resolves to something simple: AI competition moved from "whose single GPU is faster" to "whose rack-scale system is cheaper." Falling inference economics directly decide whether the agent products you use can cut prices and raise responsiveness. As an AI-using business, you need not understand NVLink, but it is worth caring which generation of infrastructure a vendor runs on, because that often explains your bill better than the model name does. The generational gap in infrastructure is the real switch on the 2026 AI cost curve.
