ArticleOctober 11, 2025Free to read

RAGFlow: a Heavyweight Knowledge Base Engine

RAGFlow offers deeper, more granular RAG settings than existing knowledge base products, including advanced features like Rerank Model, RAPTOR, and Self-RAG.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • RAGFlow is an open-source “heavyweight knowledge base engine” that offers deeper, more granular RAG settings than existing knowledge base products, including advanced features like Rerank Model, RAPTOR, and Self-RAG.
  • RAGFlow is deployed through Docker, supports all mainstream large model providers (cloud and local), and offers rich options for knowledge base creation and Assistant customization.
  • RAPTOR technology builds a tree structure through multi-level summarization, improving reasoning on complex questions; Self-RAG uses self-reflection by the large model to solve the problem of over-retrieval.
  • RAGFlow’s professionalism is reflected in its careful choice of document chunking methods and its comprehensive settings for retrieval and generation.
  • The article emphasizes that RAGFlow, as an engine, supports integration with other Chatbots or Agents through APIs, making it an ideal choice for building a local knowledge base.

I’m recommending a heavyweight product to everyone.

If you’re not satisfied with current knowledge base products and want to improve retrieval accuracy, I recommend trying RAGFlow. It provides deeper and more detailed options, so you can make targeted adjustments based on the specifics of your documents.

If your team or company wants to set up a local knowledge base, I also suggest you study RAGFlow first. I’ve said this in the community before: more likely than not, the thing you tinker together yourselves won’t be better than it—it’s better to build customizations on top of it.

RAGFlow is an open-source RAG engine. It integrates a lot of technology deeply, and it updates very quickly. Let’s compare:

Knowledge base apps like AnythingLLM generally only let you choose which embedding engine to use in RAG settings, plus how big the Chunk Size is and how much Overlap there is.

Now look at RAGFlow. In addition to the Embedding Model, you can also choose a Rerank Model. In the knowledge base settings, you can choose different chunking methods for different document types, and whether to enable RAPTOR. In the Chatbot settings, you can choose whether to enable Self-RAG.

Simply put, RAPTOR first splits a document into small chunks, then summarizes each chunk, and then summarizes again to form a higher-level abstraction. By layering these summaries on top of one another, it ultimately forms a tree-like structure. For complex questions that require multi-step reasoning, turning on RAPTOR works better.

As for Self-RAG, as the name suggests, it self-reflects. Because although RAG solves the problem of supplementing external knowledge, in actual use it can sometimes run into issues like over-retrieval. So you need the large model to judge and self-reflect.

So as you can see, these more advanced things are not in the apps we commonly use right now; they’re still lightweight designs. RAGFlow is positioned as an engine, so it has to be strong enough itself and go deep enough technically. Since it’s an engine, it has to support exporting horsepower outward. Through RESTful APIs, RAGFlow can connect with other Chatbots or Agents—I’ll introduce this in detail in the community later.

In this video, I’ll first walk everyone through deployment and usage.

Deployment is easy with Docker. The only thing to note is to leave at least 50G of disk space, because this project is pretty big.

First, create a RAGFlow folder. Then open it in VS Code and use the git clone command to clone the repository locally. Then use the cd command to enter the docker folder. Finally, run the docker compose up command, and it will start downloading the images.

Because it includes some model files, the project is quite large. You have to be patient and wait. It took me about 10 minutes, but everything finally finished. Start the project in Docker. Open a browser page and enter localhost, and you’ll see the RAGFlow page.

The first time you enter, you need to register, which is also convenient for team use. First click the avatar in the upper right corner and make some settings. I won’t go over username, avatar, password, and the like—everyone already knows those. The main thing is the model settings here.

RAGFlow supports all mainstream model providers. In China, there are Moonshot, Zhipu, and others; overseas, there’s OpenAI and the rest. Basically, it has everything you’d expect.

For cloud platforms, fill in the API Key and click confirm, and it will verify whether it can be used. Then open the dropdown list and you can see the supported models, including Chat, Embedding, Image2Text, and Speech2Text.

If you’re running locally, like with Ollama, remember that Base URL should be host.docker.internal:11434, not localhost:11434. Don’t enter the model name incorrectly. If you’re not sure, open the terminal and type Ollama list, and it will list all the models you currently have. Then copy and paste the names over.

Once the settings are complete, you can create a knowledge base. There are mainly three things here:

First, which embedding model to use. You can use the one that comes with RAGFlow, or your own.

Second, the chunking method. RAGFlow has different chunking methods for different document types. Select any one of them, and you’ll get a specific explanation on the right. You can see its professionalism from this. If you’re not sure, you can also choose General.

Third, after choosing the chunking method, you may be asked to set the chunk size. The default is 128, and you can increase it based on the document, for example to 256 or 512.

As for RAPTOR at the bottom, you can turn it on and try the effect.

Once all these settings are done, you can upload documents. I prepared an article about Nvidia here, copied from a WeChat public account, and it’s about Nvidia’s networking products.

Everyone knows that Nvidia’s GPU and CUDA are its moat. Now the trend is changing—the performance of a single card is no longer enough to meet the needs of large language model training and inference, and clusters are the only way out. But combining tens of thousands of GPUs into a super-large GPU is extremely difficult. So Nvidia is building its third moat: Networking.

I’ve gone off on a tangent; let’s get back to RAGFlow. After the document is uploaded, don’t forget to click Start manually. Sometimes parsing fails; it’s fine, just try again. If it still doesn’t work, go back and adjust the settings—maybe the chunk size is set too high, and so on.

Once it’s done, we can see all the text chunks.

To test retrieval performance, RAGFlow also provides Retrieval Testing. We can enter a question and see which relevant text chunks it finds.

For some scenarios, such as AI customer service, we want retrieval to be as accurate as possible. So we can test it at this step and go back to modify it if we’re not satisfied.

Finally comes the implementation stage. At this point, you need to create an Assistant, which is the chatbot. Likewise, RAGFlow also offers rich customization options.

For example, what opening line the AI should use to greet the user; how the AI should respond if it doesn’t retrieve relevant content from the knowledge base; whether to enable Self-RAG; which knowledge bases to associate with it, and so on.

I’m sure these three pages of settings already exceed most Chatbot products on the market. Because RAG actually contains two parts: one is Retrieval, and the other is Generation. The settings in this step are meant to improve the quality of generation, and they’re very easy for people to overlook—RAG can’t just be about retrieval.

OK, once everything is done, let’s test a question and see how it answers: Why does Nvidia need to build switches?

Although the response took a bit long, the result was pretty good. And that’s even without me making detailed settings. I believe that if I spend some time tuning it, the result will definitely be very good.

OK, that’s all for RAGFlow deployment and basic usage. From retrieval to generation, the settings it provides should meet a wide range of needs. That’s also why I called it the Ultimate RAG Engine in the community. If I have more advanced content later, I’ll post it in the community. It’s a waste to go too deep in public. See you in the next one!