ArticleOctober 11, 2025Free to read

You Can Also Run DeepSeek R1 Locally on Your Phone

DeepSeek R1 can be deployed and run on local devices like phones, and free apps like PocketPal AI support it.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • The DeepSeek R1 model can be deployed and run on local devices like phones, and free apps like PocketPal AI support it.
  • Local deployment of AI models has advantages such as stable operation, fast speed, being free, a rich selection of models, freedom in use, and privacy and data security. It is the trend in AI adoption.
  • The release of DeepSeek R1 is a major positive for the AI industry, pushing AI adoption, encouraging competition among model vendors, and sparking reflection on compute usage and the value of open-source models.
  • On desktop, Ollama is the best tool for locally deploying large models. It supports a variety of open-source models and can be combined with front-end tools like Open WebUI.
  • For mobile local deployment, the main choice is small models such as 1.5B. In the future, as technology develops, mobile AI capabilities will be even stronger.

You don’t necessarily have to use the official app to use DeepSeek R1. You can also run it locally. In fact, on a phone.

The one I have in my hand is an iPhone 12 mini. It’s already as old as it can get, and I’ve never wanted to replace it. Yet it can actually run R1, which surprised me a lot.

I’m using PocketPal AI, a free app I recommended in the community before. I downloaded the 1.5B model file with Q4 precision, and generation is pretty smooth. As you can see, just like in the official app, it first gives the reasoning process and then the result. In tests on the Benchmark page, you can see detailed numbers: about 20 tokens per second; peak memory usage is about 33%.

If you have a newer iPhone, you can download higher precision and get better results. For example, I tested it on my wife’s iPhone 14. The highest it could run was Q8 precision, at 16 tokens per second. Anything higher had no response, like FP16.

To be honest, compared with DeepSeek R1 1.5B, I personally prefer Qwen2.5 1.5B. R1’s reasoning process is too verbose, and the final result doesn’t necessarily offer a qualitative improvement. Anyway, everyone should choose based on their own situation and preferences. Right now there still isn’t a model that significantly outperforms all the others. And I also think a good model isn’t necessarily suitable for you.

Also, I know that after this video goes out, people will definitely question the necessity of local deployment again. Every time I post this kind of video, I get criticized. So let me give a unified response here.

Old netizens should remember that many years ago, Google launched the Chromebook, a netbook. Its office software was all web-based apps, the full Google suite. By those people’s logic, wasn’t that enough? Why still need a local version of the Office suite? In the end, the market gave its answer.

The same is true for AI deployment on-device. If everything relies on cloud compute, AI absolutely cannot become widespread. For example, it requires network access; when too many people use it, there may be queues; and there are also bizarre cases of it becoming dumber and lazier. All of these limit our use of AI. In addition, there are privacy and data security issues.

So, relying on on-device compute, running 1.5B or 3B models on mobile, and 7B or 14B models on desktop, is definitely the development trend for the next one or two years.

For super-individuals, having more compute means being able to run more powerful models. Knowing how to use AI on each kind of device means you can access AI more freely. Put together, these things can give you an Unfair Advantage over ordinary people.

Hello everyone, welcome to my channel. Modestly speaking, I’m one of the few bloggers in China who can clearly explain the Why and How of AI. What I provide is more valuable than tutorials. Remember to follow. If you want to connect with me, come to the newtype community. More than 800 friends have already joined for a fee!

Back to today’s topic: deploying DeepSeek R1 on-device.

The New Year holiday period was especially lively. Before the holiday, Trump launched a coin. It seemed very unreasonable, but if you think about it carefully, it wasn’t a big deal. If someone is going to smash everything to pieces, what’s wrong with launching a coin?

Not long after that wave passed, DeepSeek arrived and made a stir throughout the whole holiday. My view is simple: this is a major positive for everyone.

First, a free and open-source model that supports deep thinking and web search, and has the strongest Chinese-language ability, can let more ordinary users use AI.

I saw in my social feed that many friends who basically didn’t use AI before started using DeepSeek this time. A few days ago, when I was dining with relatives, an aunt actually brought up DeepSeek on her own and recommended their app to me, insisting that I download it and try it.

Making AI widely accessible is truly a meritorious thing.

Second, after R1 was released, people in the industry started reflecting in all kinds of ways. For example, whether previous use of compute was too wasteful, and so on. At the same time, it also gave closed-source vendors more urgency, such as OpenAI, to quickly release new models and products. See, didn’t O3 mini come out?

I believe that after this wave, every model vendor can gain something. That is the meaning of open source and open weights. When some people used to say, “open source is a tax on stupidity” and “open-source models will only get further and further behind,” doesn’t that look especially ridiculous now?

Third, for investors, this wave was both an opportunity to sell Nvidia and an opportunity to buy Nvidia. On the day of the big drop, I started buying. The logic is simple, and I also posted it in the community:

If DeepSeek’s method is scalable, then buying chips still has to continue.

They did not discover, from zero to one, a new path different from the Scaling Law. It is still the original big direction. And there is also no situation where CUDA is not needed, high compute is not needed, GPUs are not needed, and ASICs are used instead. All of that is just outsiders pretending to understand and fooling you for traffic. Companies will still try every possible way to buy chips, such as through Singapore.

So this drop was just temporary panic, and because prices had already risen so much before, the market broadly expected a pullback and was waiting for a new story. So everyone more or less acted out this scene together:

The general public was happy and felt vindicated. Capital cashed out and started waiting and watching. The U.S. government also had a reason to call for stricter controls. Everyone got what they wanted. We all have a bright future.

I still firmly believe that in AI, there is no such thing as overtaking on a bend.

Chinese people are especially good at doing things from 1 to 100. This was especially obvious in the internet and mobile internet eras. Because the basic R&D work from zero to one had already been done by others and shared out. Then we followed up and did application implementation. Look at China’s VCs—what firm really dares to invest in true zero-to-one projects? The investment results they brag about are all harvesting existing dividends.

But this wave of AI is different—basic R&D and implementation applications are advancing side by side. So it won’t work to just wait for the fruit to ripen and pick it. Nobody wants to be the sucker either.

DeepSeek and domestic AI companies are very different, whether in money or in talent. Maybe that’s the reason they succeeded.

Alright, I can’t say too much more about this topic. I’ll post a video in the community later to explain it in detail. Let’s get back to talking about deploying DeepSeek R1 on-device.

For everyday use, if you’re on desktop, the simplest method is definitely through our old friend—Ollama.

Go to the DeepSeek R1 page on the Ollama official website, and you’ll see the original model as well as six distilled small models, ranging from 1.5B and 7B to 70B. I tested both a PC with a 3060 GPU and an M4 Mac mini.

On the 3060, running 7B gives 46 tokens per second, very smooth. Running 8B gives 44 tokens per second, about the same. Running 14B drops to 26, which is still completely acceptable.

Note: this is data with OBS screen recording on. Without it, the tokens per second would be four or five higher.

Now let’s look at the M4 Mac mini. With 24GB unified memory, running 7B gives 19 tokens per second. Running 8B gives 17 tokens per second. Running 14B drops to just 10.

It seems that the Mac mini’s main advantage is power consumption. If you’re pursuing performance, a PC is still the way to go.

Once the model is running, there are many app options for chatting with it.

If you don’t need so many features and just want something cleaner, you can use Enchanted.

If you also need RAG and similar features, you can use AnythingLLM, which I recommended many times last year. It’s very easy to install and doesn’t require Docker. Docker really turns a lot of people away.

In addition, products like LobeChat and Typingmind all support connecting to Ollama. The ecosystem here is already very rich, and everyone can pick freely.

If you want to use it on mobile, 7B definitely won’t run, so you can only choose a 1.5B size.

As for the apps needed to run the model, there are also quite a few choices. For example, I previously paid for this one. Its advantage is that, besides supporting local running, it can also connect to OpenRouter or your own server. But its drawback is that it supports too few open-source models, only the ones in the list.

So I ultimately chose PocketPal AI. It supports downloading model files from Hugging Face. That feeling is like connecting to a vast ocean.

Open the app and tap the plus button in the bottom-right corner. At this point, you can choose to load from local storage, meaning a model file you’ve already downloaded, or download from Hugging Face. I chose Hugging Face here. Just enter a few keywords in the input box and you can find the model you want.

After that, using it is very simple: load the model and start chatting. The only thing to pay attention to is to increase the context length in settings. Otherwise, you may only get the reasoning process and not the final result.

Open-source model development is moving very fast today. New models generally cover the full range of sizes. For example, Alibaba’s Qwen 2.5 includes seven sizes; VL also has a 3B version. Imagine this: in another half a year to a year, the small sizes that can still run on phones will have stronger performance and more mature multimodality. By then, you’ll understand the benefits and necessity of local deployment.

OK, that’s all for this episode. If you want to discuss AI further, come to our newtype community. I’m there. See you next time!