ArticleOctober 10, 2025Free to read

How to Build an Agent System

I made a small upgrade to my note-taking system. Before, the system was just an “offline version,” only able to generate new content based on what was already there. After the upgrade, it became an “online version.”

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • Agent is the core of AI agents, used to automate task execution, and the key to building one lies in clearly defining requirements and workflow design.
  • A Multi-Agent System uses role-based collaboration to handle complex tasks, such as a combination of Researcher, Editor, and Note Taker.
  • In addition to the LLM as the “brain,” an Agent also needs tools as its “hands and feet,” such as a search tool (Tavily) and a note tool (Obsidian).
  • Building an Agent system requires Python scripts. Even if your programming ability isn’t strong, you can still modify and assemble existing scripts.
  • RAG and Agent are key technologies for AI-native applications, and understanding and practicing them can improve AI usage efficiency.

Full Content

I made a small upgrade to my note-taking system.

Before, the system was just an “offline version,” only able to generate new content based on what was already there.

After the upgrade, the system became an “online version”: I added AI search and report-generation functions. And once everything is done, it also automatically generates a note, so I don’t have to manually paste it into Obsidian anymore.

Behind these functions is the power of Agent / AI agents.

In my previous video, I introduced the basic concept of Agent. Some friends said they wanted to see a concrete example. So this episode is also a simple demonstration, to show you how an Agent is built and how it works.

Although there are now quite a few tools, such as difi.ai and the like, that let you set things up with just a few clicks, if you want to fully realize your own needs and do everything exactly the way you want, you still have to rely on code.

But there’s no need to worry. First, there are lots of ready-made Python scripts online; you can just make a few small changes and piece them together, and they’ll work. Second, it doesn’t require you to have a very high level of programming skill. Being able to understand it is enough. You could even treat it like English reading comprehension for CET-4. Someone at my elementary-school level can get started with it, so you definitely won’t have a problem.

OK, let’s get to the point.

Agents are for getting work done. So the starting point for everything is definitely the requirements, and the clearer they are, the better.

My requirement is very simple, and it comes from situations I often run into in daily life:

When I’m organizing notes or writing in Obsidian, I often need to look up some information. After searching through a bunch of web pages, I need to create a new note, extract the useful content from them, organize it, turn it into something with a logical structure, and save it in my notes so it’s convenient for the next step.

I want to hand over these tedious, low-value technical tasks to several Agents working together.

As I said in newtype on Knowledge Planet, the most important thing in building a Multi-Agent System is: how do you want it to work?

So, to meet this requirement, we need three roles, each responsible for one task:

Researcher: responsible for searching the web for information, then summarizing what it finds into a report. Editor: strong content ability and good writing skills; responsible for writing a note based on the report provided by the Researcher. Note Taker: its job is very simple: create a new note in Obsidian, then paste in what the Editor wrote.

This is a very simple division of labor, and it’s easy to understand. The difficulty lies in what tools to equip the Agent with.

You can think of the LLM as a separate brain, like the kind in sci-fi movies. It only has the ability to “think,” not the ability to act. So besides giving the Agent the LLM brain, we also have to give it tools—we can’t have it go off and work empty-handed, right?

Based on the division of labor, the Agent needs two tools:

Search tool: with this, the Agent can search the internet. Note tool: the Agent needs to know where the notes are stored, what format to use, and what the title of the new note should be called.

There are already many search tools available today. For example, Google and DuckduckGO can both be used directly. I chose Tavily. The search API they provide is specifically optimized for LLMs and RAG, and it works quite well. You can use it by adding just two lines of code.

For the note tool, you need to use a bit of ingenuity, because Obsidian doesn’t provide an interface for other programs to connect in and create notes. But there is still a solution:

All of Obsidian’s notes are in md format. So we can just create an md-format file directly in the folder where the notes are stored. In other words, we bypass the step of creating a note inside the software by creating it externally.

So, based on this solution, we end up with these few lines of CustomTools code, which specify the location of the note folder and the naming rule for the file—namely, naming it according to the time the note was created.

After putting all of this together, it forms a script like this, which includes these parts:

Basic settings, including what the API Key is, which specific model to use, and the tool settings. The three Agents I just introduced, what each of them is responsible for, and what tools they are allowed to use. Several subtasks to be completed, and which Agents participate in each subtask.

Once all of this is assembled, you run the script, wait ten-odd seconds, and the task is done.

In the future, each time I use it, I only need to modify this one line—that is, tell the Agent what I want it to search for.

Actually, I could also use Gradio to add a visual interface. But since I’m the only one using it, I don’t need to be too particular.

Following the same logic, we can make some modifications to this script. For example, we can input a public-account article link, let the Agent read it, pull out all the content, distill and summarize it, and finally save it into the notes.

What I’m introducing here are all the simplest workflows. The main point is to give everyone a concept. If you really want to build a bigger project, the whole system design becomes much more complicated. It will use more tools and more LLMs, and the collaboration among Agents, as well as between Agents and users, will become more complex too.

OK, that’s all for this episode. I hope that through this episode and the previous one, everyone can develop a basic understanding of Agents. As I said before: RAG and Agent are the key to using AI well. If you have any questions, come find me on newtype on Knowledge Planet. See you next time!