ArticleOctober 11, 2025Free to read

Let AI Pretend to Think: Thinking-Claude

Although it is “performance,” this approach can force the model to produce more logical and more comprehensive answers, which has practical value.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • Thinking-Claude uses prompts to let AI “pretend to think,” showing its thought process and improving answer quality.
  • The output of a large language model is completed in one shot; its “thinking” is performance rather than a real feedback loop.
  • Although it is “performance,” this approach can force the model to produce more logical and more comprehensive answers, which has practical value.
  • The article uses Apple’s M chip unifying desktop and mobile SoCs as an example to show how Thinking-Claude provides a valuable reasoning process.
  • The project is implemented through a browser plugin and prompt settings, making it simple to use and worth trying.

Recently, a 17-year-old teenager blew up overnight. His name is Tu Jinhao. He used a set of prompts to guide Claude on how to think more deeply and more systematically, hoping to get higher-quality answers in the end.

Not only that, Tu Jinhao also made a browser plugin, with the goal of increasing credibility. Because once it’s installed, we can see the AI’s specific thought process.

Driven by domestic self-media, within less than 24 hours, little Tu was elevated to the throne. Some people even used this to mock OpenAI—if “slow thinking” can be achieved with prompts alone, then what exactly were you all spending so much effort on with the o1 model?

To be honest, when I first saw this project, I was also quite surprised. Because based on my experience, prompt engineering doesn’t have that much power; otherwise, we wouldn’t have needed to do Multi-Agent before to discuss things and correct errors. But after I started using it, it really did seem like Claude thought first and then gave the answer.

So, where is the problem?

Actually, Thinking-Claude is playing a “magic trick” — it makes the AI pretend to think, to perform thinking. The thought process we see has already been generated; it’s all scripted.

Hello everyone, welcome to my channel. Modestly speaking, I’m one of the few bloggers in China who can explain the Why and How of AI clearly. What I provide is more valuable than tutorials. Remember to follow me. If you want to connect with me, come to the newtype community. There are already 600 friends who have paid to join!

Back to today’s topic: Thinking-Claude.

The reason I say the thinking brought by Thinking-Claude is fake and is a kind of performance is that the output of a large language model is completed in one shot. It is trained to generate continuously, predicting the next token based on the previous one. Even if you force it to pause halfway through generating tokens, it still cannot judge the tokens it has already generated. Because it simply doesn’t have that mechanism by nature, how could a few of your prompts suddenly make it enlightened?

To put it in a simple analogy, a large language model is like a faucet. Once it’s turned on, it gushes out according to a preset path. And this mechanism is at odds with thinking.

Think about how we write. During the whole writing process, we pause, we go back and revise, and sometimes we even start over from scratch. That is thinking. That is what thinking looks like—it is not single-threaded, but a feedback-loop mechanism.

A large language model doesn’t have this mechanism; it can only tell you the whole answer in one breath. If it wants to improve, it has to wait for the next round of conversation. But now it is being forced to think. What should it do?

It can only perform for you. Just like a child performing for adults.

The large language model splits its output into two parts: one part serves as the answer, and the other part serves as the thinking.

This is a limitation of technology and architecture, and prompt engineering alone can’t break through it. So don’t ever think OpenAI was a sucker, spending so much money training an o1 only to lose to a 17-year-old middle school student. These self-media people are way too folk-science.

But even though the thinking is fake, even though it is a performance, I still keep using it, and I recommend that everyone use it too. Because performance has its own value.

You can think of it this way: before installing this project, Claude is improvising. Fortunately, it has a good foundation and strong acting skills, so no matter how it acts, it’s not too bad. It’s just that sometimes the lines are a little off, after all they’re being blurted out on the spot, with no polish.

Using Thinking-Claude is like stuffing Claude with a script and forcing it to act according to the script. So there is definitely improvement. It’s just that sometimes the improvement is greater, and sometimes it’s smaller.

That is why I say “performance has its own value.” Under the constraints of this performance framework, forcing the model to produce more logical and more comprehensive answers—why not?

And honestly, I quite like the thinking it performs; it often gives me inspiration and is very valuable.

For example, I asked Claude: why does Apple need to unify desktop and mobile device SoCs through the M chip?

If you just look at the answer, it’s pretty good, but it seems like something is missing. Let’s look at the thinking it performed.

At the beginning, Claude thought this was a technical issue. Because x86 and ARM are two completely different instruction set architectures, developers need to maintain two sets of code, which adds a lot of repetitive work and cost.

Let me add a little more explanation here. In plain language, if you want a machine to understand your code, you have to translate it—translate your code into language the machine can understand. This process is called “compilation,” compiling source code into machine code for a specific architecture.

x86 and ARM correspond to two completely different machine languages. That means developers have to do the translation twice, and then check whether there are any problems, as well as make necessary optimizations—this is truly a very heavy burden. So Apple needed to unify the architecture.

Moreover, once everything switched to its own chips, it no longer had to coordinate with Intel. Apple could do whatever it wanted, on its own schedule. This kind of strategic freedom is very appealing.

Immediately after that, Claude’s thinking expanded from the strategic angle to the business angle and the user experience angle. Adopting a unified architecture means lower processor costs, more room for profit, and business benefits. And iOS apps can run directly on Mac, which also improves the user experience.

Overall, a unified architecture is not just a technical choice; it is a strategic decision at the ecosystem level. A closer, more unified ecosystem can give Apple greater dominance in market competition and technological innovation. This is the kind of thing a company of Apple’s level would do, and should do.

Did you notice? The thinking Claude performs actually gives us a reasoning process. I think this value is no lower than the formal answer.

So, if you can install it, install it and give it a try. The process is very simple—just two parts, anyone with hands can do it.

First, install the browser plugin.

Go to the GitHub page for Thinking-Claude, download the entire project as a compressed package, and unzip it. Go to your browser’s extension management page, turn on Developer Mode, and then load the extention folder in the project.

Second, set the prompt.

Go to the GitHub page again, open this file, and copy everything inside. Then go to Claude, create a new Project or open an existing one. On the right, you can add prompts. Paste everything you just copied there, and you’re done.

Perhaps sensing the danger of domestic self-media hype, in a later update, Tu Jinhao specifically added a passage emphasizing that Claude’s abilities are all pre-designed and cannot achieve a massive improvement. And this project is only meant to show us the model’s “inner monologue.”

OK, that’s it for this episode. If you want to talk about AI, come to our newtype community. See you next time!