Original video in Chinese.
Key Takeaway
- A large language model is treated as an “operating system” above all operating systems, with components like memory management (context length), a file system (chat history, knowledge base), drivers (Function Call), and a user interface (natural language interaction).
- OpenAI is upgrading GPT and ChatGPT according to operating-system logic, such as increasing context length, Function Call capability, memory functions, and the speed of natural language interaction.
- The “operating system-ization” of large language models will cause them to “eat up” many application categories, and for entrepreneurs, room for survival will be squeezed.
- Using the phidata project as an example, this article shows how an Agent, RAG, and GPT-4o can be assembled into a simple operating system.
Why do all the internet giants want to do large language models?
Because large language models are the operating system above all operating systems.
You think your product experience is good enough, but no matter how good it is, the user still has to do it themselves. Large language models will let most users experience, for the first time, the thrill of telling someone else what to do.
You think your technical moat is deep enough, but no matter how deep it is, it’s still only two-dimensional. In front of a higher-dimensional, sky-soaring technology like a large language model, moats and boundaries on the ground are especially laughable.
A large language model is the One Ring from The Lord of the Rings: One ring rules all.
Since it’s an operating system, it has to have the components an operating system should have.
First, memory management. For a large language model, that means context length. In the mainstream, memory capacity has already evolved from the earliest KB era to MB, and then to the GB era represented by DDR. And large language model context lengths are also improving rapidly; now 200K is thrown around all the time.
Second, the file system. For a large language model, the file system has two parts: one is the chat history. Without this, the large language model can’t remember you, and it can’t become your personal assistant. The other is the knowledge base, which everyone understands.
Third, the drivers. For a computer, drivers are used to control hardware devices. For a large language model, the driver is Function Call, function calling, which lets the large language model connect with existing operating systems, various software, and online services.
Fourth, the user interface. From the earliest command-line interaction to the later graphical interaction, they were all interaction designs based on the keyboard and mouse. Then a large language model comes along and flips the table: natural language interaction is enough, and it can even read the room. Compared with text input, through voice and facial expressions, a large language model can obtain much richer information.
What I just talked about are theories I summarized myself, and I shared them previously in Knowledge Planet newtype. And I found that OpenAI thinks the same way I do—they are updating and upgrading GPT and ChatGPT according to operating-system logic.
There’s no need to say much about context length. From GPT-3.5 to GPT-4 Turbo, from 4K, 16K, 32K, to 128K, in daily use there’s basically no need to worry about length anymore.
There’s no need to say much about Function Call either; GPT-4 is ahead by a long shot in this area.
In terms of chat history, the memory feature released in February lets ChatGPT remember the things the user wants it to remember, such as personal preferences and so on.
In terms of natural language interaction, everyone has seen the latest GPT-4o, and the response speed is already very fast. It’s said it can respond to audio input within 0.23 seconds, close to the human level.
You see, OpenAI’s ambition is at the operating-system level. And GPT-4o landing on the iPhone will definitely be a milestone event.
It’s not just OpenAI that has this same idea. I believe this will definitely become an industry consensus and direction of action. Some developers are already doing this, such as phidata. They’ve put Agent, RAG, and GPT-4o together and turned it into a simple operating system.
You can feed GPT whatever content you want to add, such as web pages or PDF documents.
You can ask GPT about any latest event, and it can go online to help you search.
You can let GPT be your investment advisor and have it help you analyze whether Nvidia stock is still worth buying.
If you want to experience this project, it’s very simple—anyone can do it.
Step one: download the compressed package containing all the files and unzip it.
Step two: create a virtual environment. For example, you can use conda to create and activate one, done in two lines of code.
Step three: install the required Librarys. Be sure to install exactly according to this txt file; don’t mess around on your own, or version conflicts will prevent it from running.
Step four: provide the OpenAI and EXA API Keys to the system through this export command.
Step five: open docker and install PgVector.
Step six: use Streamlit to turn this code into an APP and run it. Open a local link, and you can see the interface and features demonstrated just now.
These functions, a few months ago, were all separate projects one by one. For example, RAG was RAG, and Agent was Agent. In the past month, I’ve noticed everyone suddenly starting to build integrations.
Behind this, there is both technological progress and an iteration in everyone’s understanding. You can tell from the content in my Knowledge Planet:
At the beginning, everyone was asking me about local large language models and knowledge bases. Now people are asking about Agents more and more. The overall level, everyone’s level, is improving.
And I have a feeling, or rather a rough judgment:
Since a large language model is a highly centralized operating system, it will definitely eat up many, many application categories. For entrepreneurs, maybe you can only wait until this monster has eaten almost enough before you can get a piece of the pie.
So don’t rush to make a move.
OK, that’s all for this episode. See you next time!