Original video in Chinese.
Key Takeaway
- Meta’s open-source Llama 3.1 is a milestone; its performance reaches GPT-4o level, and it can be adapted to specific tasks and domains through knowledge distillation and fine-tuning.
- Fine-tuning is the process of training a general-purpose large model, like a college graduate, into one that masters specific skills, like company training.
- LoRa and QLoRa are fine-tuning techniques that efficiently modify a model by adding “sticky notes” instead of rewriting the whole model.
- A dataset is the “training material” for fine-tuning. The Alpaca dataset helps train high-quality instruction-following AI assistants.
- SFTTrainer is a training tool that simplifies the fine-tuning process, like a training teacher.
- Overfitting means the model “studies by rote” and loses the ability to apply what it has learned to new situations; it needs to be avoided through parameter settings like lora dropout.
- The key to fine-tuning a large model lies in the “quality of the textbook” (the dataset) and the “quality of the teaching” (parameter settings).
- The Unsloth framework is a magic tool for accelerating fine-tuning, significantly reducing VRAM usage and training time, making it easier for beginners to use.
- A fine-tuned model can be exported as an adapter (skill module) or a GGUF file, and uploaded to platforms like Hugging Face.
Full Content
Meta’s open-source Llama 3.1 is a tremendously valuable thing.
Because the best closed-source model represents the ceiling, the upper limit of what humans can achieve. And the best open-source model represents public welfare, the baseline that everyone can access, a manifestation of the value of technological equality.
This time, the open-source Llama 3.1 reaches GPT-4o level in performance. We can use knowledge distillation to use the biggest and strongest 405B model to build smaller models; we can also use fine-tuning to adapt an 8B model to specific tasks and domains.
Before, some people in China said that open-source models are a fool’s tax, that open-source models will only fall further and further behind. People like that are either stupid or bad, nothing more than clowns.
OK, enough digression. Today I want to talk about fine-tuning. I hadn’t touched this field before because I felt the conditions weren’t there yet. Now the models are strong enough, and the tools are mature too. I tried it, and it was much easier than I expected—you see, last week I posted in the newtype community saying I was going to fine-tune Llama 3.1 with Unsloth, and by the afternoon it had already succeeded.
I ran the whole process on Google Colab, using a free T4 GPU. The dataset wasn’t large, and training took 11 and a half minutes. Generating the three GGUF files, q4, q5, and q8, was slower; I waited for probably more than half an hour. In the end, these GGUF files were all automatically uploaded to my Hugging Face account.
The reason it went so quickly and smoothly is mainly because I used the Unsloth framework. This framework really is a magic tool for accelerating fine-tuning. After using it, VRAM usage is lower, and training time is also significantly shorter. I strongly recommend everyone try it.
To make it easier for beginners like me to use, Unsloth provides both the model and the code. I only made a few modifications based on what they provided.
Although there isn’t much that needs to be done by hand in the whole process, you still need to understand the knowledge related to fine-tuning, because there are deep tricks involved. I’ll first share, in plain language, some of the points I think are most important, and then I’ll walk everyone through the code, otherwise it’ll just be a blur.
First, what is fine-tuning?
When a vendor trains a large model, it’s like a college student graduating successfully, already equipped with a certain level of general skills. But to get hired and take up a job, they still need company training. That company training is fine-tuning, which lets the large model, as a newcomer, quickly master some specific skills.
Second, what are lora and qlora?
If we compare a large model to an encyclopedia, then when we do fine-tuning, we are not rewriting the whole book. We’re just sticking sticky notes on some pages with extra information written on them. LoRa is that kind of sticky note. QLoRa goes a step further: its sticky notes can write more words on smaller pieces of paper.
Third, what is a dataset?
As I said earlier, a large model needs “on-the-job training.” So the dataset is the training material. You can turn your own data into dataset format, or use public datasets. Among public datasets, in order to help large models better understand human instructions and respond appropriately, researchers at Stanford University created the Alpaca dataset. With it, we can train high-quality instruction-following AI assistants with relatively few resources.
Fourth, what is SFTTrainer?
For us users, SFTTrainer is a training tool. It simplifies the fine-tuning process and provides many configuration and optimization options, which is especially useful. For the large model, SFTTrainer is like the teacher at a training class. It receives those large-model students, takes the dataset as the textbook, and then starts teaching the large model how to better perform specific tasks.
Fifth, what is overfitting?
We’ve all met people who study too hard: they do great on exams, but if they encounter a problem the textbook didn’t cover, they can’t handle it. The same possibility exists for large models: they may only be able to deal with situations they’ve seen before, and lose the ability to apply what they’ve learned to new cases. The result of this kind of “rote learning” is called overfitting.
So, based on these five points, we can conclude that fine-tuning a large model has two keys:
First, the quality of the textbook. If the dataset is no good, then no amount of training will help.
Second, the quality of the teaching. How to use limited resources to teach the large model just right involves a lot of parameter settings, and that’s where the real tricks are.
Next, I’ll show everyone the code I used for my first fine-tuning last week. Don’t feel intimidated; this is just a process of getting familiar with it. After going through it a few times, these code lines will feel much friendlier. It’s actually very simple. After we finish looking at it, you’ll know that the most core settings are the “teaching settings” and the “textbook settings.”
At the very beginning, of course, I install and load all the required packages.
Then I load the model that Unsloth has already preprocessed. Mainstream models are all included, such as Mistral and Gemma. My target is Llama 3.1, so I just fill in Llama 3.1 in the model name field. Unsloth’s Hugging Face homepage has more models, including ones like qwen, so everyone can take a look.
In this setup, there is a parameter called max seq length. It means the maximum number of tokens the model can process at one time. Different models have different default values, from 512, 1024, 2048, and even more. You can simply understand it as: if a large model is reading a textbook, how many characters of content it can look at in one go.
Once this step is done, the next thing is parameter configuration. Among them, target modules refers to which specific part of the model we plan to modify. If we compare a large model to a robot, this robot already knows some basic movements. At this point, we want to teach it how to dance, so we modify the movement modules in its legs instead of changing the whole robot. Once this is set, the entire fine-tuning process becomes more targeted and more efficient.
In addition, there are two other important parameters:
The larger the lora alpha value is set, the more significant the influence of lora becomes. In other words, we can use this setting to balance the model’s original performance and the new skills.
lora dropout means that during training, a certain proportion of neurons will be randomly turned off. It’s like when you practice the piano, and sometimes you play with your eyes closed. That forces you to improvise freely and prevents “rote learning,” or in other words, overfitting, from happening.
Once the model is configured, the next step is to configure the dataset. My goal is to strengthen Llama 3.1’s Python capabilities, so the textbook I prepared for it is python code instructions. The content format of this dataset includes three columns:
Instruction is the instruction given, Input is the specific input, and Output is the ideal result the model should produce.
Following this format, we repeatedly train the large model so that it knows what kind of feedback to give when it encounters this kind of instruction and this kind of input. It’s just like when we used to do practice problems.
When we get to the training stage, there is a max steps value we need to consider. This is like when you’re doing bench presses at the gym: one set is 12 reps, and you stop when you reach that number. But this value has to be set just right. Because if it’s set too high, it may lead to overfitting or waste computing resources; if it’s set too low, the large model may not have finished learning before you stop it.
As you can see, apart from importing the dataset as the textbook, all of the above settings are closely related to teaching quality. Teaching students is very tricky; it’s not something you can just hand to any teacher and be done with it. The same goes for teaching large models. It also requires targeted configurations based on different needs, different models, and different datasets.
There isn’t much to do in the actual training process further down, so we just watch. According to this dataset and my settings, this training took nearly 12 minutes and used only 68% of VRAM—Unsloth really is impressive.
Finally, after the model is trained, we need to export it. There are two simple methods:
Export only the adapter. This adapter isn’t the model; you can understand it as a skill module.
Or export GGUF files and upload them to your own Hugging Face page. Here you need to fill in your Hugging Face token, which can be generated in the backend of the website.
I chose q4, q5, and q8, so it took quite a while. After everything was done, when I went to my own page, I could see the GGUF files. The files everyone usually downloads from Hugging Face come from this process.
That’s the whole fine-tuning process. If your machine is powerful enough, you can run it locally. If you just want to try it out, you can use the free Google Colab.
After this fine-tuning succeeded, I felt a strong sense of accomplishment, so I upgraded from the free version to the Pro version. It seems to be just a few dozen yuan a month. Next I’ll do more fine-tuning, accumulate more experience, and then turn it into a video to share.
Because the content after this will be relatively niche, I’ve decided to post it only inside the community as an exclusive video. If you’re interested, then join newtype.
OK, that’s it for this episode. See you next time!