Original video in Chinese.
Key Takeaway
- The M4 Mac mini is an ideal choice for a lightweight AI server: low power draw and strong performance, enough to meet the needs of running large language models locally.
- Ollama is an excellent tool for running large language models locally. It supports multiple models and precisions, and with the right settings it can keep models resident in memory for an always-ready experience.
- By changing Ollama’s listening address, other devices on the local network, such as a phone, can access the local large language model.
- Enchanted is a clean, smooth app on iOS for connecting to Ollama, making it a great choice for mobile use.
- The article emphasizes the advantages of deploying open-source large language models locally to solve problems like unstable cloud services and high costs.
Full Content
I hereby declare that the M4 Mac mini is my personal best gadget of the year. Seriously, it’s just too good!
I bought this machine with 24G of memory and a 512G SSD on Xianyu. The shop helped me buy it in Macau, and then SF Express shipped it to Beijing. All in all, it cost 7,000 yuan. I checked the price on the mainland China official website and found that it was actually 500 yuan cheaper.
In other words, if you buy the mainland China version, you spend more money and still get a “crippled version.” How does that make any sense? I really don’t understand.
After I got the Mac mini, the first software I installed was Ollama, and then I downloaded Qwen 2.5. Because I had always wanted to make a setup like this happen:
A machine that is powerful enough and cool enough to serve as a server running local large language models, and then provide them to all devices on the local network, such as a phone.
Before this, I had always used this PC to run large language models. But the power consumption and noise were really not something I could tolerate leaving on all the time. Even though reason told me it didn’t actually consume that much, I still couldn’t feel at ease. So the M4 Mac mini finally made my idea a reality.
Now, as long as I’m at home, I can use a local large language model right from my phone. For some reason, I find this kind of self-hosting strangely satisfying. It’s a completely different experience from using someone else’s service.
Actually, I can also connect to the Mac mini at home when I’m outside by using ngrok, which I introduced before, to set up intranet tunneling. But if I do that, the speed drops, so forget it.
Hello everyone, welcome to my channel. Modestly speaking, I’m one of the few creators in China who can explain the why and how of AI clearly. What I provide is more valuable than tutorials. Remember to follow me. If you want to connect with me, come join the newtype community. More than 600 friends have already paid to join!
Back to today’s topic: running a large language model on the M4 Mac mini.
I’m planning to do an upgrade before the Spring Festival, with the goal of completely solving my day-to-day AI usage problems. Right now, whether it’s ChatGPT or Claude, using them in China always feels unreliable. For example, account bans are completely uncontrollable. Once it stops working, you’re stuck. Using the M4 Mac mini as a lightweight server to run a large language model is my first attempt.
Let’s start with a simple test and see what size of large language model this 24G unified-memory machine can run. The standard is simple: how many tokens per second it can output.
The testing tool is Ollama. Turn on Verbose, and you can see the running speed.
As for the models, I downloaded two sizes, 7b and 14b, each in two precisions, Q4 and Q8, for a total of four models. There’s no need to think about 32b; it definitely won’t run, so there’s no need to test it.
At Q4 precision, the 7b generation speed is around 20 tokens per second, which is especially smooth and fluid. And 14b is around 11 tokens per second.
My own intuition is that a speed of 11 is basically the lowest acceptable threshold; anything slower definitely won’t do. At 20, it counts as smooth.
Let’s look at the Q8 speed next. At this precision, the 7b speed drops to around 13 tokens per second. And 14b is even lower.
So, all things considered, for the M4 chip with 24G of unified memory, my personal choice is to run a 14b model at Q4 precision. The speed is acceptable to me, and the completeness of the answers is clearly better than 7b. I tried leaving it running for more than half an hour, and it was basically just warm, which made me feel pretty reassured.
OK, the model is chosen, but that’s not all — Ollama still needs a few settings.
In the default state, if it’s idle for five minutes, Ollama will automatically release the model. That means if we suddenly need it and want to chat, we have to wait for Ollama to load the model again — and that’s really annoying, right?
So the first setting we need to change is to set OLLAMA_KEEP_ALIVE to -1. That way, it won’t automatically release memory, and we can achieve the goal of being ready to respond at any time.
The second is a network setting. I learned this from asking Cursor.
In the default state, Ollama only listens on localhost. To let other devices on the local network, such as a phone, access Ollama too, we need to change its listening address.
Enter this line in the terminal: OLLAMA_HOST=“0.0.0.0:11434” ollama serve
0.0.0.0 means Ollama will listen on all network interfaces. No matter where the request comes from, it accepts it. 11434 is its default port, so there’s no need to change it. After making this change, devices like phones and iPads can access Ollama through the local network IP address.
So now the last question is: what app should I use on mobile to connect to Ollama?
On desktop there are too many choices, such as the classic Open WebUI, and a bunch of AI plugins for Obsidian all support it. On iPhone, my personal choice is Enchanted, for three reasons:
First, this app is especially simple — it’s pure chat, and it supports both text and voice. There are no random extra features, so it fits my needs perfectly.
Second, it has that smooth, native iOS feel. If you’re going to use something long term, this kind of experience matters a lot.
Third, Enchanted supports Ollama. Just fill in the address and port and you’re good to go, very convenient. Of course, the reason I didn’t choose LM Studio is also because it only supports Ollama.
Today’s open-source large language models are already strong enough. Quantized versions can meet everyday conversational needs. Paired with the M4 Mac mini, it really feels great. I strongly recommend everyone set up a system like this and try it.
OK, that’s it for this episode. If you want to talk AI, come join our newtype community. See you next time!