Original video in Chinese.
Key Takeaway
- Agent platforms fall into two camps: ecosystem-based ones, like DingTalk, and tool-and-workflow-based ones, like dify. dify creates Multi-Agent Systems by providing knowledge bases and tools.
- If you want to learn Agent, you should start with dify, because it presents code logic in the form of an intuitive flowchart, making it easier to understand and practice.
- dify’s workflow design emphasizes the integrity of logic and process; the large language model only steps in when needed, rather than dominating everything.
- Workflows can make conditional judgments and branch based on user input, enabling more fine-grained task execution.
- dify workflow examples, such as text summarization, show how to combine knowledge bases and prompts to improve a large language model’s professional capabilities.
- Practicing Agent through dify helps build a basic understanding of Multi-Agent Systems and lays the foundation for learning other Agent frameworks.
Full Content
There are two major schools of Agent platforms:
One is ecosystem-based. For example, DingTalk.
On DingTalk, a lot of a company’s business is already carried there, and a lot of internal data has accumulated there. At that point, adding an Agent on top of the existing ecosystem, so that enterprises can call on the capabilities of the large language model and build intelligent workflows around those capabilities, is a very natural thing to do.
The second is tool-and-workflow-based. For example, dify.
dify provides the two foundations needed to create a Multi-Agent System:
A knowledge base and tools. For tools, you can use existing ones, or create your own. On top of these two foundations, you can then build a chatbot, an agent, or a whole set of workflows.
After watching my previous videos, many friends came to DM me asking how they should learn Agent. My suggestion is to start with dify, which is good at tools and workflows. There are two reasons:
First, as I’ve repeatedly said in the Knowledge Planet newtype, the most core part of Agent is not technology, but the workflow—how exactly you want them to do things.
dify does this in a particularly intuitive way: it presents the logic of code on a canvas in the form of a workflow. You’ll understand it as soon as you use it. I’ll demonstrate it in a moment.
Second, and this is something I’ve always emphasized: learning by doing.
For us, AI is not a theoretical question, but a practical one. And dify is especially suitable for deconstructing and assembling. Just treat it as a toy, as building blocks. When you get a Workflow running end to end, you not only learn something, but also get a real sense of accomplishment.
So, how should you get started in practice? It’s simple:
First, see how others do it. The official dify documentation provides many ready-made workflows. Pick any one that interests you, break it open, and study it. Then build a simple one yourself with your own hands.
Let me walk everyone through an official workflow sample called the “Text Summarization Workflow.”
Generally speaking, a workflow starts with the user’s input. In this text summarization workflow, it asks the user to enter the text to be summarized, and to choose whether the summary should be an overview or a technical abstract:
If it’s just an overview, then it’s simple—just let the large language model handle it directly. If it’s a technical abstract, then it involves a lot of specialized concepts and expressions, so you need to use a knowledge base, because the large language model’s pretraining data does not include this domain knowledge.
The first step is to let the user make one of two choices. Then in the second step, you need to make a conditional judgment based on the user’s choice, using if and else—which should feel very familiar to friends with programming experience.
Because there is a conditional judgment, a branch appears in the third step, just like I said earlier:
If what the user wants involves professional content, then go retrieve from the knowledge base. Then send the text the user wants summarized, together with the relevant content found in the knowledge base, to GPT-3.5.
If the user simply wants a summary of the text, then just send the text that needs summarizing to GPT-3.5, skipping the knowledge base retrieval step. It will be a bit faster.
After the fourth step of the branching is completed, the fifth step is to merge the two branches. In any case, just take the result and send it to the sixth step, wrap it in a template, and everything is done.
This is a typical workflow. The reason I’m bringing it up is to help everyone understand the way they think:
First, the large language model is not everything; it only steps in at the parts where it needs to play a role. What matters most is logic and process, something holistic that requires you to have a global view.
Like that branch just now—if you didn’t deliberately ask the user to make a choice at the beginning, and didn’t add a conditional judgment later, then you’d have no choice but to retrieve from the knowledge base no matter what, and that would be much slower.
Second, if a knowledge base is involved, you need to provide the large language model with two things: the information retrieved from the knowledge base, and the user’s original request. This step is the same as the workflow in RAG.
These two inputs can be explained clearly in the large language model’s prompt. If you want, you can also tell the large language model the format you expect here; in fact, that is basically the expected output in CrewAI.
Besides the official sample I just demonstrated, I also recommend looking at the others, so you can see what kinds of possibilities there generally are. For example:
If you need to decide how the next step should be executed based on the user’s input, besides the if/else conditional judgment I mentioned just now, you can also use a “question classification condition”—based on different content, look for reference materials in the corresponding knowledge base, and then let the large language model answer.
Once you’ve really digested these ready-made workflows, you can assemble one yourself. Once you get it running, you’ll have a basic understanding of Multi-Agent Systems.
If later you learn an Agent framework, such as AutoGen, you’ll find that the logic is the same. And with the understanding you’ve built on dify, using an Agent framework should feel much more natural.
OK, that’s it for this episode. If there’s anything you want to talk about, come find me in the Knowledge Planet newtype—I’m always there. See you next time!