ArticleOctober 11, 2025Free to read

GraphRAG: Very Good, But Very Expensive!

GraphRAG performs excellently on complex queries, but its current GPT-4 usage is costly, and running local large language models runs into slow speeds and errors.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • GraphRAG is Microsoft’s open-source retrieval-augmented generation technology that combines knowledge graphs, aiming to improve the accuracy of AI knowledge bases and address the limitations of traditional RAG, which cannot capture complex entity relationships and hierarchical structures.
  • By extracting entities and their relationships, GraphRAG builds a massive knowledge graph, thereby achieving a “global” advantage.
  • Deploying GraphRAG requires installing the relevant libraries, creating directories, placing documents, initializing the project, configuring the API key and model parameters, and building the index.
  • GraphRAG performs excellently on complex queries, but its current GPT-4 usage is costly, and running local large language models runs into slow speeds and errors.
  • Microsoft open-sourced GraphRAG in hopes of improving its speed and cost with the help of the community, so it can be applied more widely.

Microsoft recently open-sourced GraphRAG. This is a retrieval-augmented generation technology that combines knowledge graphs. Put simply, it can significantly improve the performance of AI knowledge bases, allowing AI to answer the complex questions you ask more accurately based on the documents you provide.

In this video, I’ll talk about why GraphRAG is needed and how much it costs at this stage. Really, the cost is terrifyingly high.

When it comes to AI knowledge bases, we often hear people in our community complain: the accuracy isn’t good enough, and the AI never quite answers the point. One of the root causes of this problem is the limitations of traditional RAG.

When we use this technology to build a knowledge base, the entire indexing and retrieval process is based on text chunks. Put simply, we cut a large document into many smaller text chunks; when a request comes in, we look for which text chunks are the most relevant and best matched; finally, we give the retrieved text chunks to the large language model together with the request as reference material.

This technology has two limitations:

First, it cannot effectively capture complex relationships and hierarchical structures between entities.

Second, it usually can only retrieve a fixed number of the most relevant text chunks.

Put these two together, and traditional RAG struggles especially badly with complex queries. For example, if you give it a novel and ask, “What is the theme of this book?” there’s a good chance it won’t give a reliable answer.

To make up for the shortcomings of traditional RAG, Microsoft introduced and open-sourced GraphRAG.

As I said a few days ago in the newtype community, the core of this technology is one keyword: globality.

When GraphRAG builds an index for a dataset, it does two things:

First, it extracts entities.

Second, it extracts the relationships between entities.

Visually, these entities are dots, and two related entities are connected by a line. In this way, a massive knowledge graph takes shape—that’s where the Graph in its name comes from, and it’s the clever part of this technology.

Because to express complex relationships, one very effective approach is to process them in the form of a graph. You can think back to detective shows and crime dramas you’ve seen before—don’t they often show an entire wall of clue boards? That is actually the most intuitive way to represent complex relationships with a graph, and it means the same thing as what we’re talking about today.

Because it uses a knowledge graph, GraphRAG can grasp complex, subtle data relationships, and that’s how it can build a global advantage and improve RAG accuracy.

OK, the why is done. Let’s talk about the how, that is, how to use it.

I recommend everyone follow the official beginner tutorial and run through it once. It’s only a few lines of commands, and it went very smoothly for me on Mac, with no errors at all.

Step 1: pip install graphrag — nothing much to say here, very standard. There are quite a lot of things to download, so please wait patiently.

Step 2: create a directory called ragtest, and create a folder named input inside it.

Step 3: put a document into the folder. The sample document provided by the official docs is Charles Dickens’ A Christmas Carol. After downloading it, put it into the input folder you just created and name it book.txt.

Step 4: initialize the entire project. At this point, we’ll see several more files. The two most important ones are these:

One is the .env file, where you fill in the OpenAI API Key.

The other is settings.yaml, which is used to set the model and various parameters needed for encoding and embedding. If you want to use a local large language model, set it here. I’ll demonstrate that later.

Step 5: once everything is ready, you can create the index. This process will be relatively slow; I waited several minutes.

Step 6: you can now officially start Q&A. As mentioned earlier, GraphRAG’s strength is “globality.” So as a test, the question naturally is: “What is the theme of this story?”

After the request is submitted, we can see GraphRAG begin processing and outputting according to the configuration requirements in the settings file, such as which model to use and the maximum token count.

The final result is pretty good. You have to know, this is a novel of nearly 200 pages. If it weren’t for building a global knowledge graph, there would be no way to handle a question like this.

But everything has a cost. For just one novel, using GPT-4 to create the index and do one round of Q&A actually cost me $11!

The reason it’s so expensive is that, in order to handle this document, GraphRAG made 449 API requests to call GPT-4. By comparison, the embedding model was only called 19 times.

This price is really outrageously high. Even if it dropped to $1, it would still be expensive—I’d upload a slightly larger document and a cup of Luckin would be gone.

So, the question everyone cares about comes up: what if we switch to a local large language model?

There’s absolutely no problem on the configuration side. For example, on my PC I used LM Studio to run Llama 3 and nomic embed at the same time. In the settings file, change the API Key to lm-studio—actually it’s not needed, it’s just there to satisfy the format requirements; change the API Base to localhost:1234/v1 (if it’s Ollama, then it’s 11434); then just fill in the model name. The embedding model below is filled in the same way.

After saving, follow the same process again. This time, I ran into two problems:

First, the entity extraction process was extremely slow. I waited something like 20 minutes. When using OpenAI’s model before, it was done in just a few minutes. This difference in time should be caused by differences in model performance. After all, the scale is what it is—the Llama 3 I was running locally was only 8B, which is far behind GPT-4.

Second, after finally finishing extraction and reaching the embedding step, it kept throwing errors and simply could not move forward. I tried switching the embedding model back to OpenAI’s, but it still didn’t work; at most it got to over 70% embedding and then errored out again. I struggled with it all night and really didn’t have the energy to keep burning time on it, so I had to give up.

Actually, even if it didn’t error out, taking more than half an hour to process a large document is not acceptable in real use.

My guess is that this is why Microsoft open-sourced GraphRAG: they want to rely on the community to help optimize it. After all, with the current speed and cost, even if the generated answers are great, it’s still a money-losing proposition.

OK, that’s it for this episode. If you want to chat with me, come join the newtype community—I’m always there. See you next time!