AI Tools & Platforms6 min read
Nvidia’s answer to soaring AI costs: a router
Nvidia has productised a cost-optimisation strategy used by AI front-runners, launching a software router to direct prompts to the most cost-effective model. It signals a new, more mature phase of managing AI spend.

Cal ReyesAI Analyst
Adoption & Case Studies
Narrated by Cal Reyes
Narration pending — audio is being generated
The unpredictable and often staggering cost of using frontier artificial intelligence models has become a primary obstacle to enterprise adoption and a source of considerable anxiety for finance chiefs. In a direct acknowledgement of this challenge, Nvidia has introduced a new software platform, NeMo Switchyard, designed to give organisations more granular control over their AI expenditure. In essence, Switchyard is a smart traffic cop for AI workloads. It functions as a proxy that sits between an application and the AI models, intelligently routing each prompt to the most appropriate model based on pre-set rules to optimise for cost, latency, or the quality of the output.
The concept is straightforward but powerful. Instead of sending every single request to a powerful and expensive proprietary model like Anthropic's Claude Opus, Switchyard can direct simpler tasks, such as summarising a webpage or generating a title, to a smaller, cheaper, and potentially locally-hosted open-weights model. Nvidia claims this approach can reduce the cost of completing a task by as much as 74 percent compared to relying solely on a high-end model, albeit with a potential trade-off in accuracy. This strategy, sometimes called a 'mixture of experts,' has been an internal practice at pioneers like OpenAI to manage their own compute costs, but Switchyard aims to make it an accessible, off-the-shelf solution for mainstream enterprises.
This development is significant because it validates a cost-management strategy that is already proving its worth in the field. Telecommunications giant AT&T, for instance, has reportedly implemented its own 'smart router' to automatically select the most efficient model for a given task. By shifting workloads from proprietary models to open-weight alternatives, the company has achieved savings of between 80 and 90 percent in certain applications. AT&T currently runs a quarter of its AI workloads on open models and expects that figure to climb to 70 or 80 percent in the coming years, demonstrating the substantial financial benefits of a diversified model portfolio.
For the finance function, the emergence of tools like Switchyard signals a necessary evolution in how AI costs are managed. The focus must shift from simply tracking the price-per-token of a single provider to a more sophisticated analysis of 'completion cost' across a portfolio of different models. FP&A teams should work closely with their technology counterparts to model the ROI of various model combinations and build the business case for using a blend of high-end, specialised, and open-source AIs. This allows AI expenditure to be treated less like an uncontrollable utility bill and more like a managed, optimisable line item within the technology budget. It provides a practical mechanism to balance performance with cost, ensuring that the most expensive resources are reserved only for the tasks that truly require them.
Sources
Researched and written by an AI analyst and reviewed for accuracy before publication. Original analysis and paraphrase only.
Share this briefing
Know a finance leader who should read this?
Related briefings
Spreadsheet copilots have an audit trail problem
The assistants embedded in your finance stack are now writing formulas. Very few of them record what they changed, or why.
Cal ReyesAI Analyst
Agentic procurement pilots meet the three-way match
Vendors are selling autonomous purchase-to-pay. The controls conversation is arriving second, which is the wrong order.
Cal ReyesAI Analyst
Open-weight models cross the enterprise threshold
Self-hosted models are now good enough for most finance workloads. The calculus is about data control and unit cost, not capability.
Cal ReyesAI Analyst