Original video in Chinese.
Key Takeaway
- Running a large language model locally has advantages like stable operation, fast speed, no cost, a rich selection of models, and greater freedom of use.
- Running a large language model locally requires a certain level of hardware. The minimum recommended setup is 16GB of RAM and 4GB of VRAM, with an ideal setup being 32GB of RAM and 24GB of VRAM or more.
- Different AI tasks, such as image, audio, and text generation, have different hardware requirements.
- Tools like LM Studio can help users visually understand how well their local hardware supports large language models.
- As AI becomes more widespread, on-device AI will become the trend, and AI PCs and AI phones will hit the market.
How powerful does your computer need to be to run a large language model locally?
A lot of friends have messaged me privately asking this question. In this episode, I’ll give everyone a unified answer. But before that, there’s one question we need to answer first:
ChatGPT works fine, so why bother with these open-source large language models or open-source projects on a local computer?
It’s simple, for three reasons:
First, it runs more stably, faster, and it doesn’t cost anything.
Over the past year, between my ChatGPT Plus subscription and API usage, I should have contributed several hundred dollars to OpenAI. But I’ve been very unhappy with their servers, which often let me down.
At first, I thought it was a problem with my code, or maybe my network. Later, I tested it at a different time, and it actually worked. It was really just a flood of users from all over the world, and the servers simply couldn’t handle it—that’s the challenge of the cloud.
As AI becomes more and more popular, no matter which giant company it is, its cloud compute simply cannot keep up with this scale of demand. Moving from the cloud to local devices is definitely the trend. So this year, everyone will see more and more AI PCs and AI phones going on sale.
Because I personally judge this trend to be real, I’ve kept sharing on-device AI content in my videos and in Knowledge Planet. As a user, my actual experience is:
Running large language models and AI applications on my own computer is just so comfortable. There are no annoying issues like servers not connecting, and it’s blazing fast—that’s what natural language interaction is supposed to feel like. And I never have to feel bad about token costs anymore.
Second, there are more models and more choices.
Companies like OpenAI take the approach of building the most powerful large language model, making it general enough, and then using it to meet the needs of various vertical domains and scenarios.
But from a practical perspective, we actually don’t need models that big. For example, maybe I just want AI to help me write some code or search for some information online. There’s no need to use a cannon to kill a mosquito, especially when it also burns so much energy.
That’s where open source has the advantage. If you go take a look at Hugging Face and GitHub, it’s truly a hundred flowers blooming—there are projects of every kind. We don’t have to wait for the giants and listed companies; we can do it ourselves and be self-sufficient.
It gives me a feeling of going back to the early days of the internet.
Third, it’s especially free.
My computer usually has Ollama running, and AnythingLLM open on the frontend. Whenever I think of something, I can ask the AI anytime.
I also keep various Python scripts on hand, and when I need them, I just run them and it’s done.
Running open-source large language models locally doesn’t require the internet, but that doesn’t mean they can’t connect to the internet. I can absolutely let them access the network and bring all the material back locally for processing.
I’m currently running large language models on a desktop, and it still doesn’t feel free enough. In the second half of the year, I’ll probably get a laptop with Intel’s latest CPU. This new CPU architecture includes an NPU, which can accelerate local AI inference. I’m especially curious about how much it can actually do.
OK, that’s the benefit of running large language models locally as I see it as a heavy user. Back to today’s topic: hardware configuration.
If we break it down by use case, there are roughly these categories:
First, image generation. For example, running Stable Diffusion. The minimum configuration requires 16GB of RAM and 4GB of VRAM. I’d suggest at least 32GB of RAM and 12GB of VRAM, otherwise it’s really painful.
Second, audio generation. For example, voice cloning and music generation. At least 8GB of VRAM is needed. Ideally, with 24GB of VRAM, you can already run relatively large models.
Third, text generation. That is, various chatbots. You need at least 8GB of RAM and 4GB of VRAM. If you want to run an open-source large language model with performance similar to GPT-3.5, it’s best to prepare 32GB of RAM and 24GB of VRAM.
To make it easier for everyone to understand, let me summarize it simply:
Minimum configuration: a 3060 graphics card, 16GB of RAM. If it’s any lower than that, I really don’t recommend running large language models locally.
Ideal configuration: a 4090 graphics card, 32GB of RAM. The PC at my company has this configuration, and it’s used specifically for creative colleagues to generate all kinds of images.
As for the CPU, start with an Intel i5-12600K.
As for the PC I use for demonstrations every time, I mentioned in Knowledge Planet before that it’s all several-years-old hardware:
The CPU is an i7-9700K, and the memory is two 8GB sticks of DDR4.
At first I used the integrated graphics. After using it for a while, I found it wasn’t really workable for OBS livestreaming. To be honest, back then I was still a Bilibili gaming-category uploader, and I built that machine just for livestreaming and editing videos. In actual use, I found that GPU-based streaming was still necessary. Graphics cards were especially expensive at the time, and I could only afford a 3060.
Recently, to run large language models better, I spent less than 500 yuan to buy two more of the same memory sticks, expanding the capacity to 32GB. Finally, I can run a larger class of models.
Finally, let me tell you the most intuitive method. Go download LM Studio. This software is highly integrated, and you can directly download large language models inside it, as well as run them and chat.
When downloading, the software will give suggestions based on your machine’s configuration. For example, which models can run and which ones definitely won’t work. That way, you’ll know what to expect.
Then at the usage stage, you can drag the slider on the right to adjust how much the GPU is involved. With the default setting, it relies more on memory, so it runs a bit slowly. If you drag the slider all the way to the end and use the GPU fully, the speed increases a lot. So Nvidia really makes that much money because they have the ability to do it—there’s nothing much to say about that.
OK, that’s it for this episode, about why you should run large language models locally and the recommended hardware configuration. If everyone wants to discuss further, or if you have any questions for me, come find me in Knowledge Planet. See you next time!