When talking about Nvidia, you can’t avoid CUDA. But most people’s understanding of CUDA stops at “Nvidia’s programming language.” That isn’t wrong, but it seriously understates its weight.
CUDA is not a piece of software. It is an entire ecosystem Nvidia has built on top of GPU hardware over nearly two decades. Today, almost all AI training code running on GPUs around the world is written around CUDA. This wall is Nvidia’s real moat.
Let’s start with the facts.
In 2026, Nvidia’s share of the AI accelerator market was about 80%, AMD’s was about 7% to 8%, and the rest was made up of in-house chips from hyperscale cloud players.
AMD’s MI300X hardware specs are not bad. Its per-card FP8 compute is close to Nvidia’s same-generation products, its HBM capacity is even larger, and its price is only half as much. By rights, this should be a very good product. But its market share just won’t go up. Why?
The answer is one word: software ecosystem.
CUDA has an entire toolchain—cuDNN, cuBLAS, TensorRT, NCCL, Nsight. The underlying calls of mainstream frameworks like PyTorch, TensorFlow, and JAX are all deeply tied to CUDA. Researchers write code with “import torch,” but what actually runs underneath is CUDA; engineers deploy with TensorRT, and under the hood it is still CUDA.
The maturity of this ecosystem comes from nearly two decades of accumulation, with Nvidia continuously investing since 2007. AMD’s counterpart is called ROCm. It started ten years later, and although it is catching up fast, the common feedback in the industry is still that stability, compatibility, toolchain completeness, and CUDA are a step apart.
The cost of that gap is very real.
For a large-model team to migrate from Nvidia to AMD, in theory the code has to be rewritten once, the tuning process has to start over, the communication libraries for distributed training have to be replaced, and when bugs appear there is no community to ask. Even if AMD hardware is half the price, once you add up labor costs and time costs, most teams would rather pay more and buy Nvidia.
This is the lock-in effect of the software ecosystem—it is not that it cannot be replaced technically, but that replacing it is too costly and not worth it.
So why can AMD still get 8% of the market?
Because on the inference side, especially when hyperscale players build their own clusters, they have the engineering capability to grind through ROCm themselves in pursuit of hardware cost performance.
Some of Meta’s and Microsoft’s inference clusters are using MI300X. But on the training side, especially frontier large-model training, almost everything is on Nvidia. SemiAnalysis has a set of data showing that AMD’s actual deployment share in AI training is far below its nominal share.
Another role of CUDA is that it allows Nvidia to keep defining the standard.
Nvidia’s NVLink Fusion, launched in 2026, opens NVLink to custom chip vendors like Marvell. In essence, this is about expanding the boundaries of the CUDA ecosystem—letting other accelerators into the circle, while Nvidia still sets the rules of the game.
This is an advanced use of a moat: not keeping enemies out, but turning enemies into part of your own ecosystem.
Once you see this wall of CUDA clearly, you can understand why Nvidia commands such a rich valuation and dares to stay expensive. It is selling not just hardware, but an entire “development environment you can’t live without.” That is also why on the secondary market, the narrative that “AMD is catching up to Nvidia” gets told every six months, but every time the market knocks it back to reality through price.
Hardware specs can be copied. Twenty years of ecosystem accumulation cannot.