ArticleApril 2, 2026Free to read

The Most Important Open Source Project of 2026: AutoResearch

AutoResearch may be the most important open source project of the year. It gives agents the ability to self-optimize, letting them evaluate and iterate on their own while carrying out tasks until the job is done. That design is incredibly important for improving agent system performance!

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • Core mechanism of AutoResearch: with just three files—program.md for defining goals/rules/constraints, train.py for the agent to modify, and prepare.py for automatic evaluation—the agent can propose hypotheses, modify code, run experiments, score itself, and iterate until it meets the target, achieving true self-evolution.
  • Extremely broad applicability: originally built for machine learning, but after community modifications it can be generalized to any measurable, structured work such as content creation and marketing emails. As long as you can define clear evaluation metrics, the agent can tirelessly iterate again and again; after 30 rounds of testing marketing emails, I saw the score improve 11x.
  • A new paradigm for human-machine collaboration: humans define the boundaries and evaluation criteria, while the agent handles execution and optimization; this is the most important open source project in the AI era. I strongly recommend getting hands-on with a quantifiable task yourself and adapting early to the productivity leap of “agent self-iteration.”

The most important open source project this year is definitely AutoResearch, without a doubt.

Because it gives agents the ability to self-optimize, letting them evaluate and iterate on their own while carrying out tasks until the job is done.

Why do I say it’s the most important?

Because today’s biggest bottleneck for agents isn’t whether they’re capable enough, but whether, after execution, they know if the result is good or not.

That’s where AutoResearch is so impressive. With just three files, it solves this problem in a very simple and elegant way.

One is program.md, which stores the goals, rules, and constraints written by humans. The agent can only read this document, not modify it.

One is train.py, which is the project file that the agent can modify freely.

And finally there is prepare.py, the evaluation script. The agent can’t modify it, otherwise it would be cheating.

After startup, the agent doesn’t generate the answer all at once. Instead, it first proposes a hypothesis, then starts making changes; after each change, it runs experiments and uses the evaluation script to score itself; good changes are kept, bad ones are discarded and tried again.

In other words, AutoResearch turns self-evaluation and iteration into a standard loop. What it can do, what it can’t do, and the standard for completion are all clearly written in the document. The agent just goes and does it, figures out how to solve it on its own, until it gets it done.

AutoResearch itself is a project for machine learning. But because it’s open source and structurally simple, it has now been modified and applied to other fields, such as advertising and marketing, content creation, and more.

As long as you can define a clear, automatically measurable metric, AutoResearch can be used in any field.

That’s why I say it’s the most important open source project this year. In this episode, let’s talk about this project in detail.

Hello everyone, welcome back to my channel. Humbly speaking, I’m one of the few creators in China who can explain the why and how of AI clearly. What I provide is worth far more than tutorials. Remember to follow. If you want to connect with me, come join our newtype community. This community has been running for 700 days, and more than 2,000 friends have paid to join.

If you’re a user in China, you can join through Knowledge Planet. If you’re overseas, you can join through Substack. My first course, daily newsletter, and exclusive videos are all available inside the community.

Back to today’s topic: AutoResearch.

AutoResearch is the latest work from AI legend Andrej Karpathy. He was previously a co-founder of OpenAI and the main person in charge of Tesla Autopilot. The hugely popular concept of “Vibe Coding” was proposed by him.

After AutoResearch was released, there wasn’t much reaction in China. But after trying it myself, I was pretty shocked. As I said in the Planet:

Human work that can be easily measured and structured will be the first to face the impact of AI replacement.

Look, first I downloaded this community-modded version of AutoResearch locally. The difference between this version and the original is that it makes it more general-purpose.

What I’m going to test is very simple: let the agent help me iterate on a marketing email so it’s more personalized, more concise, and more likely to get users to take action. I’ll give it a very, very basic email version, along with a set of evaluation criteria.

After starting it in OpenCode, the agent got to work under AutoResearch’s guidance. It first made sure it understood the project requirements and rules. Then it gave the initial version of the email a score, which was the baseline score. After that, it began 30 rounds of continuous iteration.

With every iteration, the agent measured things according to the criteria I gave it, figured out what it did right, where there was room for improvement, and what to change in the next round.

Just like that, after 30 rounds, the score for this email improved by 11x!

Although this was just a very simple and rough test, I did get Karpathy’s thinking behind AutoResearch. To sum it up, there are two points:

First, don’t be afraid of constraints. For agents, constraints may actually be a good thing.

If you give it a bunch of vague metrics, it can only rely on luck. On the contrary, if you give it a set of rules and boundaries, it can keep digging deeper and optimizing within that limited range.

Scenarios that require creative divergence are the minority. A lot of our daily work actually has clear boundaries and strict constraints. It’s just that most people, for various reasons, don’t define them clearly and end up making simple things complicated.

Second, humans define. The agent executes.

In the three documents in AutoResearch, two of them need to be defined by humans. Once those two documents are in place, the agent can start optimizing repeatedly without getting tired.

I firmly believe this is the way humans and AI will collaborate now and for a long time to come. Unfortunately, many people still can’t step out of the role of executor. Sadly, these people can only be eliminated.

I sincerely recommend that after watching this video, you try the AutoResearch project yourself. Hand over those tasks in front of you that can be measured numerically and give it a try. For these key things, you really do need to get hands-on yourself to feel it.

OK, that’s it for this episode. If you want to understand AI, want to become a super individual, and want to find like-minded people, come join our newtype community. See you next time!