ArticleOctober 10, 2025Free to read

The Best Large Language Model for Knowledge Bases

Cohere offers a generation model (Command R+), an embedding model (Embed), and a reranking model (Rerank), making it especially well suited to complex RAG workflows and multi-step tool use.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • Cohere and its Command R+ model are an “industry breath of fresh air” focused on RAG and Agent workflows, and one of its founders is one of the authors of the Transformer paper.
  • Cohere offers a generation model (Command R+), an embedding model (Embed), and a reranking model (Rerank), making it especially well suited to complex RAG workflows and multi-step tool use.
  • In some respects, Command R+ reaches GPT-4-level performance, and there is a quantized version that can run locally.
  • This article introduces how to call Command R+ via API using AnythingLLM and OpenRouter, as well as the hardware requirements for local deployment.
  • It emphasizes the importance of open-source models and open-weight models, and encourages users to try great models beyond GPT.

The AI company and large language model I’m most interested in, and my favorite, are not OpenAI and their GPT, but Cohere and their Command R+.

This company isn’t well known in China—most people only know OpenAI, and even companies at Anthropic’s level rarely get much attention. But in the industry, Cohere is definitely a presence that can’t be ignored.

Don’t be fooled by how young the company’s founders are. You should know that one of them is an author of “Attention Is All You Need.” It was this paper that kicked off this wave of large language model breakthroughs.

When they first started out, they were actually preparing to target the consumer market. Later they found that consumer products were much harder than expected, so they decisively shifted to the B2B market, helping companies put large language models into real business use. Cohere currently offers three types of models:

  1. Generation models. The Command series. They support user instructions and also have conversational abilities. The latest Command R+ is especially well suited to complex RAG workflows and multi-step tool use. In some respects, its performance even reaches GPT-4 level.
  2. Embedding models. The Embed series. Among them are multilingual embedding models, and the long list includes Chinese.
  3. Reranking models. The Rerank series. They rerank text chunks by relevance, which is key to improving retrieval accuracy.

Put simply, Cohere’s area of specialization happens to be exactly the area I’ve long focused on: RAG and Agent workflows.

I previously made many videos about personal knowledge bases because I had one judgment:

Of the two most important technologies today, crypto solves the problem of production relations, while AI solves the problem of productivity. So the application of large language model technology will definitely first land at the level of productivity tools, driven by RAG and Agent workflows.

For a long time, only a small number of companies have been willing to optimize large language models specifically for RAG and Agent use cases—most are still head-down on general-purpose large language models. So when I learned that there was still an “industry breath of fresh air” like Cohere, I kept a close eye on them.

Cohere’s latest batch of models came out some time ago. I checked recently, and the tools I use every day, which I’ve also recommended before, now support calling their API. And Command R+ now has a quantized version that can run locally. So this video came about.

Let’s start with API usage.

If you use AnythingLLM, remember to check the version number in the upper right corner. If the version number is orange, that means there’s a new version. After downloading and overwriting the installation, you’ll be able to see support for Cohere in the model dropdown.

As for the Copilot AI plugin for Obsidian, Cohere doesn’t appear in its model list, but OpenRouter does. This is a third-party platform through which you can call all kinds of large language models, including Command R+.

So all we need to do is paste the OpenRouter API Key in, then copy and paste the Command R+ name, and that’s it. After that, each time you use it, choose Vault QA for the mode and OpenRouter for the model, and you can use Command R+ to generate content.

Calling it through the API is the easiest method. If your computer is powerful enough, you can also try running it locally.

Command R+ has 104 billion parameters, so it’s a pretty large model. Even the quantized version is over 20GB. If you want to download it, you can do so through LM Studio.

My PC has 32GB of memory and a 3060 graphics card. According to LM Studio’s prompts, only three versions can run on my machine. And even if they can run, only part of the model can be loaded into VRAM. It still seems too demanding. My guess is that 64GB of memory plus a 4090 GPU should be able to run it smoothly.

Anyway, whether in the cloud or locally, I strongly recommend everyone try it. Over the past few days of using it, my impression is that Command R+ generates pretty well, and I’m very satisfied.

For future knowledge base applications, if I’m using a cloud large language model, I’ll definitely use Command R+. As for local use, I’ll still choose Qwen. It feels a bit better than the quantized version of Llama3.

One last thing: don’t just fixate on GPT as the only model. Among open-source models and open-weight models, there are many excellent ones too. Try more of them, and you may be pleasantly surprised.

OK, that’s it for this episode. See you next time!