Original video in Chinese.
Key Takeaway
- Ollama is the best tool for running open-source large language models locally, but the official models have limited support for Chinese.
- Users can download open-source large language model files in GGUF format from Hugging Face and use a Modelfile and the
ollama createcommand to add them to Ollama. - GGUF is a compressed format that allows large language models to run on consumer-grade devices, but at the cost of accuracy.
- This article explains in detail the steps for adding GGUF models to Ollama, including creating a Modelfile and using the
ollama createcommand. - It emphasizes Ollama’s openness, which allows it to run any open-source large language model and gives users more choices.
Full Content
For running open-source large language models locally, the best software right now is definitely Ollama.
However, for users in China, Ollama has one problem:
The officially provided models have very little support for Chinese.
At the moment, there is only one Chinese model, and it is fine-tuned based on Llama2. If you search for the keyword “Chinese” on the official website, you can find it (Llama2-Chinese).
If we want more options and want to use those domestic open-source large language models, what should we do? In this issue, I’ll introduce the method. Same old line:
Super simple. Anyone can do it.
In theory, aside from the models in the official model list, Ollama can run any open-source large language model as long as you have the model’s GGUF file.
So, the first question is: where do you download the model files?
The place with the most open-source large language models in the world is definitely Hugging Face. We can go to this platform and search directly for domestic large language models, such as: baichuan gguf, or qwen gguf. Through the dropdown list, you can find GGUF-format files uploaded by users.
To make the GGUF format easier to understand, you can first simply think of it as a kind of compressed format—although that explanation is not rigorous, it doesn’t really matter. Just as JPG compresses images, GGUF compresses large language models. That way, they can run on consumer-grade devices like ours.
Of course, compression comes at a cost, and that is reduced accuracy. That is why there are a series of files with different sizes in the list. Just choose a suitable one to download based on your machine’s configuration.
After downloading it, create a txt document named Modelfile, and in it you only need to write one line:
FROM D:\ollama
The purpose of this document is to tell Ollama where to find the model file. For example, in my case, it goes to drive D, the ollama folder, and finds the model file with that name.
The final step is to open the terminal and enter this command line:
ollama create
Let me explain what this command line means. Actually, you don’t need to be nervous and think that this is a command or something, and you definitely won’t understand it—you can just treat it like English reading comprehension:
ollama create, which is easy to understand, means letting Ollama create a new model file.
So what should the created model be called? That is the name that follows, which you can set however you like.
Ollama certainly can’t create a model file out of thin air, so we need to use the GGUF file we downloaded earlier. At this point, we tell it to read the txt document we just created, which contains the address of the GGUF file.
In this way, Ollama knows where to find the large language model and then create a model file with whatever name you choose.
After this command runs, wait another two or three minutes and it will be done.
In the terminal, we enter: ollama list. This command will list the models you currently have. At this point, we can see that the model we just imported already exists.
Open Open WebUI, and in the model selection dropdown list, you can also see the newest model.
OK, that’s all for the method of adding any open-source large language model to Ollama. The model version I downloaded earlier may not have great Chinese performance; it was only for demonstrating this method. Everyone can go to Hugging Face or domestic model communities to download various GGUF-format Chinese large language models, and then find the version that suits you best.
That’s all for this issue. If you have questions, or want to discuss further, come find me in Knowledge Planet. See you next time!