ArticleOctober 10, 2025Free to read

AI Needs “Shadow Clones”

Large language models have a “single-core flaw” (Degeneration-of-Thought), and multi-Agent collaboration can effectively solve complex reasoning problems.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • ChatGPT “going naked” is not enough to meet productivity needs; deploying Agents can significantly improve efficiency.
  • GPT Researcher is an out-of-the-box Agent solution, good at information gathering and report generation, and it is very cheap.
  • CrewAI is a flexible Agent framework that lets you build an Agent system freely by setting Agents, Tools, and Tasks.
  • Large language models have a “single-core flaw” (Degeneration-of-Thought), and multi-Agent collaboration can effectively solve complex reasoning problems.
  • Agent technology is developing rapidly with the support of large language models, and there will be more tools and applications in the future.

I’m not going to keep subscribing to ChatGPT anymore.

It’s fine for light use. But if you really want to use it as a productivity tool for the long term, it still isn’t quite up to it.

Let’s make a comparison. The same question:

After adding online search to GPT-4, ChatGPT gave this answer:

Pretty good, right? Let me show you what an Agent generated:

I wouldn’t say the gap is huge. It’s basically the difference between usable and unusable.

So, from a practical standpoint, I suggest everyone stop using ChatGPT “naked.” Spend a little time deploying an Agent setup, and it can save you a lot of time.

Let me introduce the two setups I’m currently using.

GPT Researcher: out of the box

GPT Researcher is a project on GitHub, mainly designed to meet the needs of information gathering and report generation — a daily work essential that really can save a lot of time.

GPT Researcher sets up two kinds of Agents:

The Planner Agent is responsible for breaking down the requirements and generating as comprehensive a set of questions as possible. After the Execution Agent gets the questions, it finds the corresponding web pages, crawls the content, and then hands it back to the Planner Agent. The latter filters and summarizes all the materials and completes the research report.

This project does two things especially well:

  1. It mixes GPT-3.5 and GPT-4 to improve speed and reduce cost. Generally speaking, one run takes about 3 minutes and costs $0.1 — that’s really dirt cheap.
  2. The Agents generated according to the requirements are domain-specific. For example, if the requirement is to do research in the financial field, then the generated Agent is a finance expert.

You only need to know a little bit of code to use GPT Researcher. Follow the GitHub tutorial: clone the repository locally, then copy and paste step by step and run the corresponding commands. If it prompts that some package is missing along the way, just install it with pip install. Finally, open a local web page and you can use it.

CrewAI: build freely

If your needs go beyond generating research reports, then you need to use an existing framework and build an Agent system yourself.

The Agent framework I’m currently using is called “CrewAI.” It looks a lot like Microsoft’s AutoGen, but once you start using it, you’ll find that CrewAI is simpler and more intuitive logically than AutoGen.

In CrewAI, you only need to set three elements:

  1. Who.
  2. What to use.
  3. What to do.

“Who” refers to the Agent. How many Agents there are, what roles they play in collaboration, what their work objectives are, what their backgrounds are, and what model they use as their brain.

“What to use” refers to Tools. The most common ones are search tools. You need to assign tools to the specific Agent that will use them.

“What to do” refers to Tasks. A project can be broken down into many tasks. Each task needs a specific description, as well as a designation of which Agents will complete it.

Once you understand this logic, CrewAI becomes extremely simple to configure.

Using report generation as an example, this is the Agent workflow I designed:

I deliberately arranged two Agents at the beginning for the requirement analysis and solution design stages. Doing this uses more tokens and takes more time, but it’s very necessary. Everything is about solving one core problem:

Large language models are especially prone to getting stuck when doing complex reasoning.

The single-core flaw

To strengthen the reasoning ability of large language models, researchers have come up with many methods. For example, the well-known Chain-of-Thought, and Self-Reflection.

But no matter how much buff you stack onto a large language model, this problem still doesn’t go away. In papers, it’s called “Degeneration-of-Thought”:

When a large language model is confident in its own answer, even if that answer is wrong, it can no longer generate new ideas through self-reflection.

Just like people, it gets immersed in its own world — mysteriously self-confident and unwilling to change.

There are many causes for this problem. For example, during pretraining, biased input concepts or flawed ways of thinking can both lead to cognitive bias.

Some problems can be solved technically, and some don’t need to be. For this one, human society actually already has a solution, and it’s the one we’re all most familiar with:

Discussion and collaboration.

No matter how smart someone is, or how high their cognitive level is, they will have blind spots.

If someone points it out — actually, sometimes you don’t even need to be pointed out; just chatting a few words with someone other than yourself is enough to climb out of it.

That’s why, even when the underlying driver is the same large language model, “multi-core” is much more reliable than “single-core.”

2024 Agent

Agents did not rise with large language models. Long before this wave of AI explosion, Agent had already been studied for many years. Large language models act as the strongest brain and solve the reasoning bottleneck for Agents, which suddenly brought Agents to everyone’s attention.

In terms of designing and deploying Agents, after AutoGen came CrewAI, and in 2024 there will definitely be more teams wanting to give it a try.

At the tool level, the number of tools that Agents can directly call is still not that large. But I believe this situation will definitely be solved this year.

This new wave driven by large language models has crowded into the frontmost ecological niche, pushing all services and products one step back. If they want to survive, they can only accept the fate of being turned into tools and being called.

As for OpenAI, maybe in one update this year, it will make targeted enhancements for Agents at the ChatGPT and API Assistant levels.

Agents are a general capability, and there’s no reason for the giants to let that slip away.