Original video in Chinese.
Key Takeaway
- Ollama is the best tool for running open-source large language models locally. It supports multiple platforms and is easy to install and use.
- Open WebUI provides a ChatGPT-style web interface, supports local large language model interaction and RAG capabilities, and can handle web pages and documents.
- Anything LLM is a more advanced local knowledge base management tool that supports multiple large language models, embedding models, and vector databases, and provides the concept of a Workspace as well as chat/query modes.
- Deploying large language models and knowledge bases locally can achieve data security, privacy protection, and more flexible customization.
- The article emphasizes Ollama’s server mode, which allows it to open a port for other software to call on its large language model capabilities.
Full Content
For running open-source large language models locally, the best software right now is definitely Ollama.
Whether you use a PC, Mac, or even a Raspberry Pi, you can run models of all sizes through Ollama. And it’s extremely extensible.
I’m going to cover how to use Ollama in detail over a few installments. In this one, I’ll start with three things:
- How to use Ollama to run a large language model locally.
- While running a local large language model, use a Web UI like ChatGPT.
- Build a fully local knowledge base.
If you have better suggestions, or if you run into any problems during installation or use, you can come find me in the newtype Knowledge Planet.
Ollama
Installing Ollama is super easy. Go to the official website ollama.com or .ai and download the corresponding version.
After installation, type ollama run in the terminal, followed by the name of the large language model you want to run. For example: ollama run llama2. The system will then automatically download the corresponding large language model files.
If you’re not sure of the model name, you can find all currently supported large language models on the model subpage of the official website. Each large language model has different versions. Choose the corresponding version based on your needs and your machine’s memory size, then just copy the command.
In general, a 7b model requires at least 8 GB of memory, 13b requires 16 GB, and 70b requires 64 GB. Use whatever you can handle; otherwise it really gets very laggy.
By default, you need to interact with the large language model in the terminal. But this approach is really too old-fashioned. What we definitely want is to operate it in a modern graphical interface. That’s when Open WebUI comes in.
Open WebUI
To install Open WebUI, you need to install Docker first.
You can simply think of Docker as a virtual container. All applications and dependencies are packaged into a container, and then run on the system.
Once Docker is ready, copy this line of command from GitHub into the terminal and execute it. If everything goes smoothly, open a local link and you’ll see a very familiar interface.
Besides basic chat functionality, this WebUI also includes RAG capabilities. Whether it’s web pages or documents, they can all be provided to the large language model as reference material.
If you want the large language model to read web page content, just add # in front of the link.
If you want the large language model to read documents, you can import them at the chat box location, or import them on the dedicated Documents page.
If you type # in the chat box, all imported documents will appear. You can select one, or simply let the large language model use all documents as reference material.
If your requirements aren’t too high, then this is enough. If you want more control over the knowledge base, download this software: Anything LLM.
Anything LLM
Ollama actually has two modes:
- Chat mode
- Server mode
What’s called server mode can be simply understood as Ollama running the large language model in the backend and then opening a port for other software to use, so those applications can call on the large language model’s capabilities.
Turning on server mode is very simple. Type two words in the terminal: ollama serve.
After it starts, paste this default link into Anything LLM. At that point, the software will read the models that can be loaded through the link. These models are the ones used to generate content.
In addition, building a knowledge base involves two other key pieces:
- Embedding Model. It is responsible for converting high-dimensional data into a lower-dimensional embedding space. This data processing step is very important in RAG.
- Vector Store. A vector database specifically used to efficiently handle large-scale vector data.
We use the defaults for both of these. In this way, the entire system runs on your computer. Of course, you can also choose to run everything in the cloud—for example, use OpenAI for both the large language model and the embedding model, and use Pinecone for the vector database. That works too.
After completing the three most basic settings, you can enter the main interface. I really like the logic of this software. It has the concept of a Workspace. Within each Workspace, you can create various chat windows and import various documents.
So you can create Workspaces based on projects—one Workspace for each project. Then import all the documents and all the web pages related to that project into the Workspace. Finally, there are also two chat modes you can set:
- Conversation mode: the large language model will answer by combining the documents you provide with the knowledge it already has.
- Query mode: the large language model will simply answer based on the documents.
This is the more advanced part I mentioned earlier about Anything LLM compared with Open WebUI. It can completely meet personal knowledge base needs. I’ve already made it the core of my desktop Workflow. Once I finish these two video installments, I’ll make a dedicated one talking about the AI tools and workflows I’m currently using.