Original video in Chinese.
Key Takeaway
- TPU ≠ a GPU replacement: TPU is a chip Google tailored for its own trillion-scale models, with a focus on “matrix specialization + systolic arrays.” It delivers extreme energy efficiency at the cost of generality; GPU is the “king of general-purpose parallelism,” with unmatched compatibility.
- Gaps across three dimensions: hardware (TPU specialized vs GPU general-purpose), software (XLA black box + code rewrites vs mature CUDA ecosystem), and networking (TPU closed-off cabinets + data center modifications vs GPU that you can just plug in and run) mean TPUs can only serve a tiny number of giants (Google/Meta/Anthropic).
- The reality: for 99% of startups, research institutions, and small and medium-sized businesses, GPU is still the only ticket in; claiming TPUs will break NVIDIA’s moat is pure ignorance, and the era when NVIDIA should be making money has already arrived.
Google has been the market’s center of attention lately. That’s because they’re the only company that has everything from chips, models, and data centers to cloud computing services, as well as consumer and enterprise applications — a true “full-stack AI” company. I’ve recommended their models and products many times.
But if you say their TPU can replace NVIDIA’s GPU, that’s just too ignorant, too naive… I completely disagree.
It’s true that Google has changed its strategy. In the past, TPUs were only for internal use; now they’ve started selling them to customers, such as Anthropic.
But the problem is that TPUs are currently only suitable for a very small number of companies. For example, companies with massive scale, long training runs for large models, and a strong talent pool that can handle the TPU stack.
For most companies and startup teams, whether you’re doing AI or research, GPU is still the only choice.
Whenever I see people blindly hyping TPUs, I think of the crowd that hyped DeepSeek earlier this year — saying things like NVIDIA’s moat had been broken and CUDA was no longer needed, etc. It’s extremely ignorant.
If you want to see clearly and avoid being fooled, you need to look at TPU and GPU from three dimensions: hardware, software, and networking. In this video, I’ll try to explain it in a way that’s as easy to understand as possible.
Hello everyone, welcome back to my channel. Humbly speaking, I’m one of the few creators in China who can clearly explain the why and how of AI. What I offer is much more valuable than tutorials. Remember to follow. If you want to connect with me, come join our newtype community. This community has been running for 600 days, and more than 1,900 friends have joined as paying members.
If you’re in China, you can join via Knowledge Planet. If you’re overseas, you can join via Substack. My first course, daily newsletter, and exclusive videos are all available inside the community.
Back to today’s topic: the difference between TPU and GPU.
Many years ago, Google, like most companies today, bought NVIDIA GPUs. One day, they did a calculation and broke out in a cold sweat. They realized that if they kept using them like this, they would become more and more dependent on NVIDIA, and in the end they’d be working for Jensen Huang. So the TPU project was launched.
At the hardware level, the core divergence between GPU and TPU is the battle between “general-purpose parallelism” and “matrix specialization.”
First, let’s look at the GPU, the “king of general-purpose computing.”
Its design philosophy is “single instruction, multiple threads” (SIMT). You can think of it as a huge army, with thousands upon thousands of small cores working in parallel, each soldier able to think independently.
To support this kind of general-purpose computing, a GPU must be equipped with extremely large and extremely fast memory, so it can frequently read and write all kinds of data.
To maintain flexibility, a GPU is designed as a standalone component: you can buy one and plug it into a computer, or buy ten thousand and put them in servers.
But everything comes at a price. If you want generality, flexibility, and all-around capability, then you can’t drive transistor utilization to the absolute limit, which inevitably means huge heat and very high power consumption.
In sharp contrast to GPUs is the TPU, the “special-purpose computing commando.”
The TPU’s design philosophy is called a “systolic array.” It doesn’t pursue individual combat; instead, it’s like a precision assembly line, with data flowing through the chip like blood, and everything serving matrix multiplication and nothing else.
This extreme specialization determines that a TPU is hard to exist as a single card. Its design unit is usually a Pod (cluster), and its memory usage is relatively restrained — because it relies on extremely high-bandwidth interconnects between chips, allowing data to be moved directly within the “assembly line,” compensating for the shortcomings of a single card.
In the end, by cutting away every circuit unrelated to AI, the TPU achieves astonishing energy efficiency. When handling specific tasks, it is cooler and more power-efficient than a GPU.
If hardware sets the ceiling, software determines how much of that potential you can actually realize. At the software level, GPU and TPU are still completely different.
The GPU’s moat is CUDA. More specifically, it’s the freedom CUDA brings.
CUDA’s philosophy is “pass-through.” It allows developers to bypass the operating system and directly command every tiny core on the GPU. You can precisely control how every thread is scheduled and how memory is allocated.
The maturity of this ecosystem is terrifyingly high. Whether you’re training large models or doing scientific computing, as long as you know how to use this toolset, you can squeeze every last drop of performance out of the GPU.
Google’s TPU is the exact opposite; it relies on XLA, accelerated linear algebra.
Because the TPU’s hardware architecture is so special, it’s very hard for humans to manually arrange how data should move through that complex pipeline. So Google chose “abstraction.”
When using a TPU, you don’t need to tell the chip, “Do step one here, step two there.” You only need to define your mathematical formula, and then hand the rest over to the XLA compiler. It’s like a super butler that automatically analyzes everything, slices up the data, lines it up, and feeds it into the TPU array.
The result of this difference is:
CUDA gives developers a sense of security, because everything is under control, and when something goes wrong, you know where to fix it.
TPU is harder to say. If the compiler “guesses” your intent correctly, performance takes off. But if it throws an error, it’s a disaster.
That’s the difference at the hardware and software levels. Beyond that, we also need to look at networking, because what everyone cares about today isn’t a single card, but cluster efficiency — that is, connecting thousands upon thousands of cards together.
And once we need to connect that many cards, the networking architecture differences between GPU and TPU fully expose the fundamental difference in their business models.
GPU networking is very much like building with blocks.
NVIDIA provides standardized interfaces (NVLink) and general-purpose network protocols (InfiniBand). Even though the blueprint is drawn by NVIDIA, the blocks are modular — you can buy Dell servers, pair them with a switch from some brand, and then plug in H100 cards.
This architecture is extremely flexible. You can put 8 cards in one cabinet, or 72; you can build a small network for R&D, or a large network for training. It is compatible with existing data center standards, and as long as you have the money, you can definitely build it.
TPU networking is much more complicated. It’s more like a “biological neural network.”
I just mentioned that TPU’s design is called a “systolic array.” This design doesn’t just exist inside the chip; it extends between chips as well.
TPU uses a technology called ICI (inter-chip interconnect), directly “welding” neighboring chips together with copper cables to build a huge 3D ring structure (Torus).
In this structure, there’s no concept of “network cable” or “network card.” The whole cluster is like a giant virtual chip.
To support this structure, Google even developed optical circuit switches, using mirrors to reflect light beams and adjust the connections. They completely abandoned traditional electronic switches.
Doing things this way creates physical isolation.
You see, GPU networking is open — you can move it into any standard data center.
But TPU networking is closed; it has its own physical world.
So once you choose TPU, you can’t just buy chips. You have to buy the entire cabinet, the entire cabling scheme, and even remodel your data center for it.
When you put hardware, software, and networking together and think about them as a whole, you’ll understand why I said at the beginning that TPUs are only suitable for a very, very small number of companies.
For the vast majority of companies, GPU is the only ticket in. Because what GPUs sell is “compatibility.”
Whether you’re a startup or a traditional company, buy a GPU, plug it in, download the general CUDA driver, and your code can run.
Although it consumes a lot of power and generates heat, it shields away all the dirty work and heavy lifting at the bottom through hardware generality.
TPUs are only suitable for giants. After all, Google originally designed them around its own needs, and they themselves are a giant.
TPU’s “cluster-based design” and “software black box” mean that what you buy back is not just a pile of chips, but an entire heterogeneous infrastructure that needs to be adapted all over again.
You need a huge demand. Because if you don’t have the need to train trillion-parameter models, you simply can’t fill up the TPU’s systolic array. The power savings you get may not even offset the migration cost.
You also need talent. You need a top-tier engineering team that can handle the XLA compiler and refactor the underlying code. By comparison, there are far more people who know CUDA.
That’s why only giants like Meta, Anthropic, or Google have the strength and qualification to choose TPU.
Only when compute scale reaches a certain level does TPU’s energy efficiency advantage get amplified into hundreds of millions of dollars in cost savings.
So after you understand the differences between TPU and GPU across these three dimensions, do you still think TPU can replace GPU?
NVIDIA weathered so much pressure back then to build this moat. Now it’s their time to make money. You may resent it, you may envy it, you may think it’s unfair — but why don’t you think about what you were doing all that time ago?
OK, that’s it for today’s content. If you want to understand AI, want to become a super individual, and want to find like-minded people, come join our newtype community. See you in the next episode!