Original video in Chinese.
Key Takeaway
- 95% automation is really semi-automation: many people let AI autonomously carry out tasks, yet still have to watch the whole process, inspect the results, and fix mistakes, so the end result is just replacing manual labor with being the “AI supervisor” — your attention gets completely tied up, and the cognitive load actually increases. The real standard for automation is not “how much AI has done,” but “can a human walk away without worry.”
- The key lies in handoff-ability: from early Cursor, which required constant intervention, to today’s Codex/Claude Code, which can be “let sleep and checked later,” AI is evolving toward fully managed operation. But the premise is that you must first understand AI’s capability boundaries — what it can reliably complete, what it can only assist with, and what must be judged by a human (investment trading is a classic example).
- Practical method: first clarify the boundaries through learning by doing, then force a verifiable acceptance standard (the super-workflow Skill in newtype OS requires at least 3 success criteria plus failure red lines). This is the dividing line between truly moving from “AI assistance” to “AI automation.”
After using AI, many people like to do automation, that is, let AI autonomously complete some tasks.
But to be honest, most of the automation people build is not really automation. Because they’ve only achieved 95% automation, and the most critical 5% hasn’t been solved.
For example, while AI is doing the work, you may have to keep watching it. Who knows what kind of nonsense it might come up with. After AI is done, you still have to inspect and correct it.
On the surface, AI saves you a lot of physical effort. But in reality, your attention is tied up, your cognitive load increases, and you become the AI’s supervisor.
In other words, you’re still working at the construction site; only your role has changed, and you’ve become the contractor, managing more “people” under you.
So I believe 95% automation is not automation — it’s actually still semi-automation. Real automation is not about “how much AI can do,” but about “whether a human can leave without worry”; the core metric should not be “how many steps were completed,” but “how much human attention has been freed.”
Put simply: the key to automation is not execution, but handoff-ability.
I’ve had a very deep personal experience with this handoff-ability over the past half year.
When Vibe Coding first started catching on, AI coding wasn’t that powerful yet. At that time, I used Cursor to develop my first product — Prompt House.
It looked like AI was coding at full speed. But there were always problems it couldn’t solve no matter what, and I had to make judgments and figure things out. And I still had to keep an eye on it to prevent it from messing up the UI I had already built or deleting code.
Honestly, I was excited back then, but I was also genuinely exhausted.
And today’s AI is much stronger for development. Basically, you can go to sleep and talk about it when you wake up.
You see, from Vibe Coding to OpenClaw a few months ago, AI has been moving toward fully handoff-able operation. But before fully handoff-able, fully autonomous operation is eventually achieved, I have two lessons to share on how to make use of AI’s current automation capabilities.
First, figure out AI’s capability boundaries.
These boundaries determine what AI can reliably complete, what it can only assist with, and what must be judged by a human.
For example, in investing, I currently have Codex monitor prices and key events for me every day. If there’s a price fluctuation or something major happens, it sends me a push notification.
In theory, I could have Codex handle trading for me — actually, the code for the order placement module is already there, and there’s no problem at all at the execution level. But I didn’t do that. Because I know:
Being able to do it does not mean being good at it, and does not mean being able to get it done.
Codex, Claude Code, and other agentic AI are the most advanced AI tools on Earth today, but their capabilities have boundaries.
For example, in investing, they might come up with a strategy that looks pretty good, but that only means it works mathematically or probabilistically, because I don’t have infinite money or time.
Knowing the boundaries means you won’t mythologize AI, and you won’t belittle AI either; instead, you’ll understand and use it in a realistic way.
So before letting AI automation run, you first have to figure things out yourself. And there’s only one way to do that:
Learning by doing.
Learn by doing.
Second, set acceptance criteria for AI.
Don’t just say, “Help me do it well”; make it clear: what should the output be, what metrics should it meet, and how should failures be handled.
This actually involves things at the level of Harness Engineering. I previously posted an article in the Knowledge Planet; if you’re unfamiliar, go look it up.
Different fields have different logics when it comes to acceptance criteria. In content creation, I specifically built a Skill into newtype OS before, called super-workflow. Its mechanism is:
When the user says they want to write an article or start a topic selection process, AI loads super-workflow. It forcibly requires acceptance criteria to be set in advance, including core information, content type, word count, tone, structure, and so on. There must be at least three verifiable success criteria, plus failure red lines.
Only after the acceptance criteria are set does AI begin the subsequent process, including outlining, section-by-section writing, reviewing, and so on.
It can be said that acceptance criteria are the core of the super-workflow Skill. What it needs to solve is the key question of how to determine whether the work was done correctly. Even for something like content creation, you still have to define verifiable standards.
Only after you understand AI’s capability boundaries will you know what kind of work you can hand over to it for automation. It is definitely not like those people raising lobsters, who just give AI the highest permissions and then let AI do everything.
After the work is determined, you still have to set acceptance criteria. That is the dividing line between “AI assistance” and “AI automation.”
OK, that’s it for this episode. If you want to learn about AI, want to become a super individual, and want to find like-minded people, come join our newtype community. See you next time!