ArticleAugust 15, 2026Free to read

Why Does Nvidia Want to Build Model Routing?

Model routing is a direction I’ve always paid close attention to. I previously introduced OpenRouter’s two-layer routing system. Recently, Nvidia has also joined in, releasing the open-source model routing tool NeMo Switchyard. A company that sells GPUs—why get into model routing? Their business logic is actually very simple and very down to earth: once you enter the Agentic AI stage, task complexity rises significantly, and execution costs rise with it. If AI usage costs can be reduced, enterprises can deploy more Agents.…

Originally published . English translation: . Read the Chinese original.

105-Why Does Nvidia Want to Build Model Routing?

Model routing is a direction I’ve always paid close attention to. I previously introduced OpenRouter’s two-layer routing system. Recently, Nvidia has also joined in, releasing the open-source model routing tool NeMo Switchyard.

A company that sells GPUs—why get into model routing?

Their business logic is actually very simple and very down to earth:

  • Once we enter the Agentic AI stage, task complexity rises significantly, and execution costs rise along with it.
  • If AI usage costs can be reduced, enterprises can deploy more Agents.
  • More Agents generate more inference requests.
  • More inference requests require more compute, networking, and servers.
  • Then, in the end, Nvidia can naturally sell more GPUs and complete AI infrastructure.

So, how do you reliably complete a high-value task at the lowest possible cost? That is the problem NeMo Switchyard and these model-routing tools are meant to solve.

At this point, you may ask: aren’t API costs continuously dropping? Why would execution costs be going up?

That logic only applies to the chatbot stage: the user asks a question, the model answers, one input, one output, simple.

But the situation now is that AI is no longer just answering a single question; it has to autonomously complete a whole sequence of steps.

For example, ask AI to complete a programming task. It may need to first read project files, then search related code, then design a modification plan, generate code, run tests, and finally keep adjusting based on error messages.

Even if each individual call becomes cheaper and cheaper, as long as the number of calls, the context size, and the number of retries grow faster, the cost of completing the whole task will still rise.

So when an Agent has to carry out so many steps, the most direct way to reduce costs is:

There’s no need to hand every task to the strongest, most expensive model. Simple tasks can be handled by cheaper models, and that’s enough.

That’s exactly what NeMo Switchyard does. It supports multiple routing methods:

The simplest approach is to judge task difficulty based on the request content. Simple questions go to small models, complex questions go to strong models.

Another approach is to assign models according to the Agent’s work stage. For example, use a strong model during the planning stage, a fast model during the execution stage, and upgrade again when tests fail or tool errors occur.

It also supports a “cheap first, then upgrade” strategy. The system first lets a lower-cost model try, then lets another model judge whether the answer is reliable. If the quality is not good enough, it calls a stronger model again.

So, how well does it work?

In LangChain’s tests, only 7% of requests were sent to frontier models, costs dropped by 74%, but accuracy also fell by 6 percentage points.

This shows that model routing can indeed save a lot of money, but the actual effect still depends on the task type and routing strategy. After proving feasibility, the next competition is about who can balance cost, speed, and accuracy better.

Once you understand the whole logic, let’s go back to the question at the beginning.

Nvidia’s move into model routing is not a cross-border leap; it is a continued expansion of its own boundaries.

In the chatbot era, the industry competed over whose model was stronger. In the Agent era, the question gradually becomes: how do multiple models work together to complete the entire task at the lowest cost?

When competition shifts from a “single model” to a “model system,” the routing layer becomes the new infrastructure.

Nvidia does not want to be only the lowest-level GPU supplier. It hopes to move upward from CUDA and inference engines all the way into the scheduling layer between models and applications.

It does not need to bet on which model will ultimately win. As long as all models keep generating tokens, and as long as more and more Agents need compute, Nvidia can keep selling its GPUs, networking, and complete AI infrastructure.

So what NeMo Switchyard really represents is not just another open-source tool release.

It represents Nvidia moving from “selling compute” toward “defining how compute should be used.”