ArticleOctober 10, 2025Free to read

LLM = OS

MemGPT treats large language models as operating systems, and uses hierarchical memory management (Main Context + External Context) to solve the context-length limitation.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • Context length is a key limitation for LLM applications, and it is very hard to improve.
  • MemGPT treats large language models as operating systems, and uses hierarchical memory management (Main Context + External Context) to solve the context-length limitation.
  • Main Context includes system instructions, conversational context, and working context; External Context includes event memory and factual records.
  • MemGPT can autonomously retrieve and edit contextual information, and has “awareness.”
  • MemGPT supports multiple backend models and can be integrated with AutoGen and other agent systems, which is highly significant for Multi-Agent Systems.

Context length is the first hurdle LLMs have to cross.

If it is too short, a lot of domain applications simply cannot be launched, such as medical GPT. Imagine that after 20 rounds of doctor-patient conversation, the doctor no longer remembers the patient’s basic situation. How is that supposed to work?

So, context length is roughly equivalent to an LLM’s memory, and it is one of the basic metrics for measuring LLM capability.

But improving an LLM’s context length is very difficult.

First, on the training side. It requires greater compute power and GPU memory, and also more long-sequence data.

Second, on the inference side. The self-attention mechanism of Transformer models requires calculating the relevance of every element in a sequence to every other element. This mechanism naturally determines that context length cannot be too long. So people have proposed a series of solutions for handling long sequences, which is another huge topic and I won’t go into it here.

MemGPT found a genius solution.

LLM = OS

What is a large language model?

MemGPT believes that, at its core, a large language model is an operating system. So context is memory, and context-length management is memory management.

How does an operating system manage memory?

Through hierarchy. CPU cache (L1, L2, and L3) is closest to the core, the fastest, but has the smallest capacity. Extending outward from there is memory, and finally the hard drive.

Depending on the need, the operating system allocates data across these three layers: the most urgent data goes into CPU cache; data that is not needed for the moment goes to the hard drive.

Since an LLM is an operating system, using the same memory-management method makes perfect sense.

That’s exactly what MemGPT does.

Main Context + External Context

This is MemGPT’s operating logic:

When an event occurs, the event information enters the virtual “memory” (Virtual Context) through the parser (Parser).

The LLM, as the processor (Processor), invokes and confirms the data in memory, and then outputs it through the parser, turning it into an action.

The key point is the Virtual Context. It is divided into two parts:

  1. Main Context: this is the context with the original length limit. Main Context consists of three parts:

  2. System Instructions. Simply put, this is the “you are a helpful assistant” that we write in the system message every time. This part is read-only and is invoked every time because it is a base-level setting.

  3. Conversational Context. It follows a “first in, first out” (FIFO) rule — after a certain length, the oldest conversations are discarded.

  4. Working Context. Simply put, this is the LLM’s notebook, which records current points of attention.

The image below fully explains what Working Context is.

When the user mentioned the two key pieces of information, “today’s birthday” and “favorite chocolate lava cake,” the LLM quickly wrote these two points into the notebook and then applied them in its reply.

  1. External Context: this is contextual information stored externally, such as on the hard drive. External Context consists of two parts:

  2. Recall Storage can be simply understood as memory of events that have already happened, and it is uncompressed, full-process memory.

  3. Archival Storage can be simply understood as a record of facts.

Once the structure is separated like this, MemGPT can autonomously operate on context.

Autonomy + Awareness

MemGPT performs two kinds of operations on contextual information:

  1. Retrieval. This needs no explanation; everyone gets it.
  2. Editing. For example, if the latest information provided by the user conflicts with previously stored information, it is edited and updated.

So contextual information flows back and forth between Main Context and External Context across the two layers.

MemGPT’s memory-management behavior is all carried out autonomously. The development team prepared a series of prompts in advance, and different situations trigger the corresponding prompt. A very important prerequisite for MemGPT to manage things successfully is:

Awareness.

It knows that it itself (the LLM) has a context limitation, so it can devise countermeasures based on the limitation (such as 16K or 32K).

You see, if AI is silicon-based life, then it starts off far stronger than humans, these carbon-based life forms, because it has clear self-awareness and self-perception, while 99.99999% of humans are in a state of “ignorance.”

In the long evolutionary process that follows, the advantages of self-awareness and self-perception will be amplified exponentially. In the end, the two species will be worlds apart. I’m digressing.

Usage + Integration

It’s very easy to try MemGPT. Open the terminal:

  1. Install with pip install pymemgpt.
  2. Set up with memgpt quickstart. If you don’t want to tinker, use memgpt quickstart –backend openai to call the OpenAI API, and it will ask you to enter a key.
  3. Start with memgpt run. If you have created an agent before, it will ask you to choose one. If you want to exit, remember to /save first and then /exit, just like playing a game — save before quitting.

MemGPT supports multiple backends. Besides the OpenAI API, it also supports various ways of calling open-source LLMs.

Most importantly, MemGPT can be integrated with AutoGen. Multi-Agent Systems are probably the most direct beneficiaries of expanded context length. As for how to implement it, the official documentation has detailed code, so I won’t show it here.

MemGPT is the kind of project that really catches my eye; it is very inspiring. If I discover more later, I’ll share it with everyone.