ArticleOctober 11, 2025Free to read

My View of AI in 2025

The keyword for AI in 2025 is Agent, and its essence is a “task engine,” not a simple “intelligent agent.”

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • The keyword for AI in 2025 is Agent, and its essence is a “task engine,” not a simple “intelligent agent.”
  • AI development will move from the “information engine” phase (led by large language models) into the “task engine” phase (led by Agents).
  • Chatbots are only the most basic form of Agent, and may be phased out in the future because they lack “context” information, which limits task-completion ability.
  • Giants that have users’ “context” information (such as Google and Apple) have a natural advantage in Agent development.
  • The product form of Agents will evolve from a “man-made form” (software/app packaging) to a “self-generated form” (AI automatically generating Agents).
  • RAG and Agents are the foundational technologies of AI-native applications, and understanding them is key to grasping the AI era.

In 2025, AI has only one keyword: Agent. Whether you’re working on models or applications, everyone will be concentrating their firepower on this point.

The stage of simply competing on models is already over. Note, I’m not saying models are no longer important. Model capabilities will definitely keep improving. But relying solely on models to fight for market share stopped working long ago. If you look back at what the three giants have done over the past six months, you’ll see it:

Anthropic added the Artifacts feature to Claude, rolled out computer-control capabilities, and introduced the MCP protocol;

OpenAI also added a canvas-like feature to ChatGPT, and put search on it, although it didn’t do a very good job;

Google used to be pretty weak, but its latest update brought it straight up to par. Gemini’s multimodality, ultra-long context, and Deep Search feature are all extremely impressive.

All the moves these three companies have made point to the same thing and mean the same thing: the model is the application. And that application is Agent.

The Agent concept has been hyped domestically for probably a year. I’ve looked around, and it seems like no one has really explained the basic logic clearly; they’ve just stayed at the vague label of “intelligent agent.” I suggest everyone forget those definitions and just remember these four words:

Task engine.

From a fundamental logic standpoint, AI is in one half a continuation of the logic of the internet.

The internet is essentially about organizing and distributing information. From the ancient era of Yahoo, Google, and various portals, to later Taobao and today’s Douyin, all of them were about reorganizing and redistributing all kinds of information, and then redrawing territories.

AI is the same. Large language models can handle information at scales that were previously unimaginable, extracting valuable knowledge and patterns from it. And beyond traditional forms like text, images, and video, AI can also handle more complex and abstract information, such as knowledge graphs, semantic networks, and so on.

So AI continues the underlying logic of the internet, continuing to organize and distribute information, and doing it even better. But that is not AI’s true mission. Agent is AI’s true face.

Information organization and distribution emphasize the static side of information. What Agent does is dynamically apply information and use it to complete specific tasks.

That’s why I think “task engine” is a more accurate and easier-to-understand description of Agent. I especially hate those three characters, “intelligent agent.” Domestic media and vendors especially love making up concepts that sound grand but say nothing.

To make it easier for everyone to understand, I’ll refine and summarize it again:

This round of AI development—namely the first phase that began with GPT-3.5—is a phase led by large language models, characterized by an “information engine.” It is a more powerful “information engine” than any product from the previous internet era or mobile internet era.

Starting in 2025, we will enter the second phase, led by Agents, characterized by a “task engine.” Agents and large language models are not separate. It is precisely because there are sufficiently powerful large language models, and precisely because there is a sufficiently powerful “information engine,” that the “task engine” becomes possible.

OK, once you understand Agent and the underlying logic of AI development, the next question is: what does Agent look like? Or rather, what is its product form?

Software and apps are product forms we are very familiar with. In the AI era, will chatbots like ChatGPT be the standard form of Agent?

I don’t think so. Chatbots are only the most, most basic form of Agent, and even this form is very likely to be phased out.

Just think about one question: for an Agent to complete tasks well, what matters most?

It’s like a person: to complete a task assigned by a boss, is the key factor personal ability? No. The key factor is “background information,” or in other words, “context.”

What are the causes and consequences of this task? What are the expectations behind the boss assigning it? What is the implication between the lines? If you don’t figure these out, what good is being highly capable?

It’s the same with an Agent. So what if your generation ability is strong? When there’s a real need, you still have to explain a huge amount first. For example, if I want to write an article, I have to tell the AI: the client’s needs are like this, the reference materials are these, and so on. And 99% of people can’t figure it out or explain it at all. The natural-language interaction we keep emphasizing today actually only suits a minority of people.

It’s precisely these prerequisites that limit our use of Chatbots. Just look at the data for these products now—how many daily active users they have, how many times they’re used per day—and you can clearly see the problem.

So a form like ChatGPT is like Mobile Monternet back in the day. People who didn’t live through that era definitely have never heard of it. In the early days of the mobile internet, Mobile Monternet was a giant supermarket that covered all kinds of information services, including SMS, MMS, mobile internet access, which was WAP, and Baibao Xiang, meaning mobile games. Doesn’t that sound a lot like ChatGPT today?

And as we all know, what truly drove the explosion and popularization of the mobile internet was product forms like Toutiao and Douyin that relied on algorithmic recommendations. If AI is going to explode and become widespread, it likewise needs this kind of “foolproof product” that suits the general public. The most critical thing among them is to make up for the “context” I just mentioned.

This is something OpenAI simply doesn’t have by nature. Who does? Google does, Apple does, Meta does, Tencent does, Alibaba does, ByteDance does.

For example, imagine this: the Chrome browser and Gemini are fully integrated. Chrome already has my saved bookmarks and all my browsing history, right? Those can serve as extremely valuable contextual information, allowing an AI version of Chrome to give me what I actually want.

That’s why I say a chatbot like ChatGPT is only the most, most basic form of Agent, and why it is very likely to be phased out. OpenAI’s current lead is only temporary. Just like Mobile Monternet back then—who still remembers it now?

OK, once you understand that “context” is the key to Agent, let’s look at product form again. I think Agent will have two forms, corresponding to two stages of development.

The first form is the current “man-made form.”

ChatGPT is an Agent, Perplexity is an Agent, Cursor is an Agent. All the Agents we have now are man-made. We package the Agent inside the shell of software or an app, and thereby complete specific tasks such as search and programming.

The number of man-made Agents won’t be too large, and they’re only a feature of the early stage. My estimate is that by at most 2026, we’ll enter the second stage and see the second form: the “self-generated form.”

As the name suggests, in the “self-generated form,” AI will automatically generate Agents. Because every person’s every need is actually wildly different. If you insist on using software or app forms to pre-extract the greatest common divisor and box everything in, you can only satisfy a portion of the shared needs.

Once the “context” I just mentioned is fully connected, all kinds of personalized needs can be turned into tasks of all sizes. Starting from the task, AI can autonomously generate the corresponding Agent to handle it. That is what the full arrival of the AI era looks like.

If you’re in investment, or if you’re a developer, you can think carefully about what I’ve said. I know that making judgments and drawing conclusions publicly will definitely attract a lot of criticism. No problem—I especially welcome everyone to dig this up six months or a year later and see who was right and who was wrong.

Most of the dozens of video episodes I made over the past year were about RAG and Agent. Back then, I said these two technologies are the foundation of all applications. To handle more related information, you must use RAG; to execute various tasks, you must use Agent. And I also made an episode before about letting AI automatically generate Agents, introducing such a technology. If I remember correctly, it should have used Microsoft’s framework.

So if you’ve been paying attention and putting it into practice all along, reaching the conclusions in this episode is very natural. Looking back from today’s point in time, you suddenly realize that everything fits together, and the direction is incredibly clear.

OK, I won’t say more. As I always say, I’m one of the few bloggers in China who can clearly explain the WHY and HOW of AI. If you want to connect with me, come to our newtype community. See you in the next episode!