Original video in Chinese.
Key Takeaway
- A Q&A engine is the next form of search engines, directly providing organized content rather than web links.
- LLocalSearch is an open-source project that lets users deploy a Q&A engine locally, with web-connected search as well.
- The basic logic of LLocalSearch is: a local LLM understands the question -> converts it into search keywords -> searches for relevant materials and stores them in a local vector database -> combines the question and materials to reason and output an answer.
- Deploying LLocalSearch requires Ollama and Docker, and downloading a Function Calling model and an embedding model.
- LLocalSearch is still at an early stage, but it offers the potential for a localized Q&A engine.
Full Content
The next form of search engines is definitely the Q&A engine.
Because what we want is not web pages, but the content inside web pages.
To peel off the shell of a web page, extract part of the content inside it, and feed it back to the user — only AI can do that.
I subscribed to Perplexity at the beginning of the year. It’s the best Q&A engine right now. The annual subscription is $200, which stings a bit, but it really can completely replace Google. Still, I canceled it recently. Because the company announced that it will insert ads into search results.
I’m honestly pretty disappointed in them: I thought they would carve out a different path, but in the end they’re going back to selling ads. And this time, with AI’s help, who knows what kind of tricks they’ll come up with.
But fortunately, we’ll soon have an alternative.
LLocalSearch is an open-source project. It’s usable right now, but it’s not perfect yet. If you want to try something new, you can give it a shot.
As the name suggests, LLocalSearch lets you deploy an entire Q&A engine on your own computer.
There’s one concept I need to clarify first: running locally does not mean it can’t connect to the internet.
This open-source project uses only the compute power of my PC and the LLM I installed on my PC. But at the same time, it has the ability to connect to the internet so it can help us look things up, right? So there’s no contradiction there.
Let me show you the result first, then I’ll talk about how to install it.
On the left is what the product looks like, and on the right is the resource usage.
Because I had OBS open for recording, GPU usage is relatively high. If OBS weren’t affecting it, the main consumption would be memory.
The basic logic of LLocalSearch is:
When you ask a question, the local LLM first understands what you mean, then converts the question into a set of keywords suitable for searching.
Next, it helps you search all relevant materials on the web and puts everything it finds into a local vector database, using Chroma DB here.
Finally, it combines the question and the materials to reason and output the final answer.
Based on the previous question, you can continue asking follow-up questions.
If you don’t trust the whole process, you can click the button in the top-right corner to expand every step.
If you want to install it too, just search for this name on GitHub: LLocalSearch. I also posted the link in Knowledge Planet, so members who have already joined can grab it there.
Before installing the project, make sure you have already installed Ollama and Docker — you need Ollama to run the LLM, and Docker to run this project.
After that, use Ollama to download these two models:
One is knoopx / hermes-2-pro-mistral, which is responsible for Function Calling. You can roughly think of it as the one that calls various tools and helps you get things done.
The other is nomic-embed-text, an embedding model with a relatively large context window.
Once the software and models are downloaded and installed, you can clone the project locally. Then use the cd command to enter the project folder and run docker-compose up, and it will install automatically.
Finally, if everything goes smoothly, open the local link localhost:3000 and you can use it normally.
LLocalSearch is still pretty rough around the edges for now, but the overall framework is there. From what I can tell, the author is just one person, a German guy. If you want to support this project, you can sponsor him on GitHub. $5 a month, $15 a month — either works. If you’re a big shot and willing to sponsor $800, then the German guy can buy a new graphics card — isn’t that more meritorious than tipping those female livestreamers with rockets?
Lastly, if you haven’t used a Q&A engine before and don’t want to go through the trouble of deploying one locally, you can try domestic options like 360AI Search and Metaso AI Search. As always:
Getting it into use matters more than anything else.
OK, that’s it for this episode. If you have any questions for me, come find me in Knowledge Planet. See you next time!