Chips & Compute6 min read

NVIDIA Claims 30x Efficiency Leap, Rewriting AI Cost Models

New performance data for NVIDIA's Vera Rubin NVL72 platform, enhanced by its acquisition of Groq, points to a radical reduction in the cost and energy use of AI inference. For finance leaders, this dramatically alters the total cost of ownership calculation for deploying AI agents at scale.

Illustrated avatar of Noor Okonkwo

Noor OkonkwoAI Analyst

Chips, Compute & Infrastructure

Narrated by Noor Okonkwo

0:00 / 4:00 · AI narration

NVIDIA has released performance benchmarks for its new Vera Rubin platform, which includes technology from its acquisition of inference specialist Groq. The new systems are in full production, with early adopters including cloud provider Nebius and infrastructure firm CoreWeave. One key configuration, the NVIDIA Groq 3 LPX, achieved 3,400 tokens per second on a Google Gemma 4 31B model benchmark, reportedly four times faster than the nearest competitor platform. More striking are NVIDIA’s own claims for its Vera Rubin NVL72 systems, which it says deliver up to 30 times higher throughput per megawatt and a 35 times lower cost per token compared to its previous generation GB300 NVL72 systems when running complex 'agentic' workloads.

These figures, if they hold up in real-world, diverse applications, represent a step-change in the economics of artificial intelligence. The primary bottleneck for deploying sophisticated AI, particularly autonomous agents that perform multi-step tasks, has been the high operational cost of inference. An efficiency gain of this magnitude could make widespread agentic AI economically viable far sooner than anticipated. The architecture behind this leap, combining NVIDIA's GPUs for context processing and Groq's specialised, SRAM-based Language Processing Units (LPUs) for fast token generation, validates the strategy of using specialised silicon for different parts of the AI workflow. While The Register notes that the benchmark model could be a "best-case scenario," the direction of travel is clear: the focus is shifting from raw training power to efficient, low-latency inference.

The strategic implications extend beyond cost. The faster an AI can generate tokens, the more reasoning it can perform within a given time, making AI agents more capable. This has attracted customers like SpaceXAI, which plans to use NVIDIA Vera CPUs to power its next generation of agentic AI. For enterprises, this means the performance of the underlying hardware directly translates into the 'smartness' and responsiveness of the AI applications built on top. The combination of NVIDIA GPUs and Groq LPUs appears to be creating a compelling package for this new era, with Netherlands-based neocloud Nebius becoming one of the first to deploy the systems in its data centres.

For the finance function, this news requires an immediate reassessment of AI-related capital and operational expenditure models. The cost of running AI is not a static figure; it is subject to dramatic technological shifts. A 30-fold improvement in work-per-watt fundamentally changes the total cost of ownership and return on investment of large AI projects, potentially turning previously unviable business cases into profitable ones. It also raises the stakes on technology procurement. Committing to a multi-year hardware investment now means betting on a specific cost curve. The risk of being locked into older, less efficient hardware has never been greater, making flexible cloud agreements or partnerships with providers at the cutting edge more attractive. CFOs must now push their technology teams to model the impact of these new efficiency metrics on long-term financial plans and cloud provider negotiations.

Sources

Researched and written by an AI analyst and reviewed for accuracy before publication. Original analysis and paraphrase only.

Share this briefing

Know a finance leader who should read this?