Original video in Chinese.
Key Takeaway
- LM Studio is the first choice for running LLMs on a Mac, especially because it supports model files optimized for M-series chips, which significantly improves performance.
- LM Studio has added support for Apple’s MLX framework, which is optimized for M-series chips and can deploy and run models efficiently.
- Through a comparison demo, this article shows the speed advantage of optimized models running on a MacBook Air with an M2 chip.
- Apple has already caught up in AI, and its hardware (unified memory architecture) and the MLX framework provide strong support for AI endpoints.
- LM Studio is evolving from a backend tool into a frontend application, signaling that AI applications are about to enter a phase of all-out competition.
If you’re using an Apple computer with an M-series chip and want to run an LLM on the machine, I strongly recommend LM Studio. Because it supports model files that have been specifically optimized for M-series chips, and they run more than a little faster.
I used this MacBook Air with an M2 chip to do a simple comparison. With the same LLM and the same prompt, the left side is the optimized version, and the right side is the GGUF version we’ve commonly used before. You can clearly see with the naked eye that the left side is much faster. Looking at the token generation speed per second, the optimized model is twice as fast.
Hello everyone, welcome to my channel. Humbly speaking, I’m one of the few creators in China who can explain both the why and the how of AI clearly. Remember to hit follow. As long as one video clicks for you, you’ll have already gotten a great deal out of it. If you want to connect with me, come to the newtype community. More than 500 friends have already joined and paid!
Back to today’s topic: LM Studio.
Among tools for running LLMs locally, LM Studio and Ollama are the two most popular. In this latest update, LM Studio has added support for MLX.
This mouthful, MLX, is an open-source machine learning framework from Apple, specially optimized for M-series chips. For example, it uses a unified memory model, corresponding to the unified memory architecture. So with this framework, you can deploy and run models very efficiently.
MLX was only open-sourced last December, so it’s still very new, but with community support it has developed quickly, and mainstream models all have corresponding versions. In the latest version of LM Studio, they also specifically added labels and filtering to make it easier for Apple users to download.
If you haven’t installed LM Studio before, you can download the corresponding version from the official website. After installation, open the software. The left sidebar contains its main functional pages, including chat mode, server mode, viewing existing models, and so on. Go to the Discover page, and you can search for and download models.
Just as I said earlier, LM Studio specifically labels MLX versions of models, so everyone can easily find them in the list. By default, it recommends Staff Pick, meaning officially recommended models. If you want more, choose Hugging Face, which will list all the models.
Different quantized versions of models have different file sizes, so choose according to your configuration and needs. If the download won’t move, it’s probably a network issue, and you’ll have to figure that out on your own.
Once the model files are downloaded, we can load them in chat mode. LM Studio provides all kinds of settings, and I’m just using the defaults here.
To make this not-so-rigorous comparison, mainly to give everyone an intuitive feel, I asked AI to help me write a Python Snake game. Since I was also recording the screen, that would affect the speed a bit. But even under those circumstances, the optimized model still runs very smoothly. And then look at the standard version—the difference is just too big.
For a long time, many people have criticized Apple for being behind in AI, and domestic media keeps writing sensational pieces about it too. But if you’re really paying attention, you’ll know that Apple has absolutely already caught up. Their accumulation in hardware far exceeds that of those PC manufacturers.
I mentioned in an exclusive community video before that Apple’s self-developed models are far ahead of the Android manufacturers next door. On the desktop side, with the MLX framework, it can fully leverage the greatest advantage of the unified memory architecture:
The CPU and GPU can directly access data in shared memory without data transfer. Small-scale tasks can be handled by the CPU. When there are compute-intensive demands, then the GPU comes in.
Hardware layer, system layer, and application layer all integrated into one—that’s what an AI endpoint should look like. This is also why I chose to make a major upgrade at this point in time: replacing the mini I’ve used for so many years with an iPhone 16 Pro—I’m planning to have a friend buy a Hong Kong version for me in November, and once the M4 version of the MacBook Pro comes out, I’ll get the 16-inch top configuration. Then I’ll share a series of hands-on experiences with everyone.
Finally, one more thing. If you’ve been using AI software all along, you’ll notice that there has been a big wave of frequent updates recently. Everyone is expanding their own territory. For example, LM Studio, which we talked about today, used to be just a more backend-oriented piece of software, helping you run LLMs locally. Now, it has moved chat mode to the front and added RAG functionality. This kind of proactive move from backend to frontend will gradually become a common choice for everyone. The phase of all-out fighting among AI applications is coming.
OK, that’s it for this episode. If you want to discuss AI further, come to our newtype community. See you in the next one!