ArticleOctober 11, 2025Free to read

Training Your Own DeepSeek-R1 with 7GB of VRAM

The Unsloth framework significantly lowers the barrier to fine-tuning large language models: with just 7GB of VRAM, you can fine-tune a 1.5B model, even on a consumer PC.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • The Unsloth framework significantly lowers the barrier to fine-tuning large language models: with just 7GB of VRAM, you can fine-tune a 1.5B model, even on a consumer PC.
  • DeepSeek’s GRPO reinforcement learning algorithm can improve a model’s reasoning ability and interpretability.
  • Fine-tuning techniques can be used to build a personal AI clone and private-domain models, enabling localized, offline AI interactions.
  • High-quality datasets and hyperparameter tuning are the keys to successful fine-tuning, but they require a lot of practice.
  • The article emphasizes the potential and value of locally deploying small models on mobile devices.

Using DeepSeek’s method for fine-tuning can significantly improve a traditional model’s thinking ability.

This is the model file I trained. I’ve already uploaded it to Hugging Face, so feel free to grab it. It’s based on Qwen2.5 3B, and after fine-tuning I strengthened its math ability. In the end, I produced three versions: Q4, Q5, and Q8. Let’s compare the Q4 accuracy. I asked a classic question:

Which is larger, 9.9 or 9.11?

First, let’s look at the original model’s answer. Not only was the answer wrong, but the reasoning was completely messed up too — what does it even mean that “the decimal parts are the same, and the only difference is in the tenths place”? That’s just nonsense.

Now let’s look at the fine-tuned version. This one is normal. The integer parts are the same, so compare the decimal parts. Naturally, you get that 9.9 is larger than 9.11.

I didn’t build this stuff myself; it’s the result of Unsloth. A few days ago, they published a blog post introducing the method and providing the code. In simple terms, Unsloth does two things:

First, it lowers the fine-tuning barrier. Small models like 1.5B can be fine-tuned with just 7GB of VRAM. For 7B and 14B models, 15GB of VRAM is enough. In other words, you can fine-tune on a consumer PC. If you use cloud compute, like the T4 GPU on Google Colab that I used, it was completed smoothly in about an hour.

Second, it improves model capability. GRPO is a reinforcement learning algorithm invented and open-sourced by DeepSeek. Using this algorithm and dataset, you can train models with stronger reasoning ability and better interpretability. Now Unsloth has applied it to fine-tuning, and the possibilities open up immediately. For example:

Private-domain model.

A business creator has their own methodology and lots of delivery cases. They organize everything they’ve accumulated over time into a dataset, including questions, answers, and solution steps. Then they use Unsloth to fine-tune a 3B model. In the end, they hand the model file to their users, whether free or paid.

After users get it, they can use the method I introduced in my previous previous episode on their phones. That means users can talk to this creator’s AI clone anytime, anywhere, without needing to be online.

For creators, in the past, only when you posted videos, wrote articles, or spoke in groups could your fans and users receive your information. Now, with this method, they can be influenced by your IP without limit.

When I previously shared methods for running models on phones, a bunch of people mocked me, saying it was meaningless and worthless. To put it bluntly: if your vision is too narrow, you deserve not to make money.

Hello everyone, welcome to my channel. Modestly speaking, I’m one of the few creators in China who can explain the why and how of AI clearly. What I provide is more valuable than tutorials. Remember to follow me. If you want to connect with me, come join the newtype community. More than 800 friends have already paid to join!

Back to today’s topic: using reinforcement learning algorithms to fine-tune models.

Before introducing Unsloth’s tools, I still need to explain the basic concepts in a way that’s easy to understand. It may not be super rigorous, but you’ll definitely get it.

In the past, reinforcement learning required a large amount of high-quality data containing solution steps, as well as a very precise, absolute reward function. Then, by brute force, you’d train the model into existence.

Later, DeepSeek found that the cost didn’t have to be so high, and it didn’t have to be such a pain — the reward function could be made more flexible. For each question, it lets the model generate a set of answers. Then it looks at that set and rewards the answer that is relatively better.

The traditional method is more like the rote teaching we used to get in school, relying on memorization to grind through problems and trying to muddle through. But with that approach, you know the what without knowing the why, so in the end you’re still trash. DeepSeek’s method, on the other hand, repeatedly thinks through the solution steps, and in the end it not only knows the what but also the why. Then the model “has an epiphany,” and a top student is born.

If that still doesn’t make sense, let me give another analogy. Traditional dog training requires clearly defining every action and designing rewards for each one. Only when the dog completes the action exactly as instructed can it get a reward.

DeepSeek’s method, however, is to have the dog perform one action three times. Among the three attempts, the relatively better one gets the reward. Then the process is repeated over and over.

Anyone who has raised a dog knows that with this kind of training method, the owner is relaxed, the dog is happy, and the results are good too.

After DeepSeek generously shared it, Unsloth picked it up and used it. But before using it, there are some limitations I need to make clear to everyone:

The model you use for fine-tuning can’t be too small. It needs to be at least 1.5B, otherwise it can’t generate the thinking markers correctly. That’s why I chose a 3B size: it meets the training requirements and can also run on a phone. Also, you need at least 300 steps before the reward really starts to increase. To get good results, I recommend training for at least 12 hours.

In the official example, the dataset used is GSM8K. It contains 8,500 high-quality elementary school math word problems. Each question takes 2 to 8 steps to solve. And the solution methods in this dataset are written in natural language rather than pure mathematical expressions. So training with it can improve the model’s multi-step mathematical reasoning.

There are several datasets similar to GSM8K, such as the MATH Dataset and MathQA. I recommend that everyone not rush to import their own dataset right away; you can practice with these first. Because after changing datasets, due to differences in format and characteristics, the reward function may need corresponding adjustments.

In addition, hyperparameter tuning also requires a lot of practice. For example:

Learning rate, which controls how fast the model learns. Set it too high, and the model may learn too quickly and miss the optimal solution; set it too low, and the model may learn too slowly and waste time.

Batch size refers to the amount of data fed to the model each time. Set it too large, and you may run out of memory; set it too small, and the model’s learning may become unstable.

Fine-tuning, like RAG, looks simple, but if you really want good results, it takes a lot of tuning. And this is not something you can be taught directly — you can only “learn by doing.” But having a barrier is a good thing. Once you cross it, you can leave a bunch of people behind.

So I bought some compute units on Google Colab, and during this period I’ll do all kinds of tests. As for the dataset, I suddenly thought that over the past year I answered a lot, a lot of questions in the Knowledge Planet. These questions can all be transformed, for example by letting the model help me batch-process them, and then putting them into a dataset.

The idea of building an AI clone and training a private-domain model through fine-tuning already came to me when I made the Llamafile video last year. Now the possibilities are getting bigger and bigger. When I have progress, I’ll talk about it in the community.

OK, that’s it for this episode. If you want to understand AI, come join our newtype community. See you next time!