Original video in Chinese.
Key Takeaway
- Perplexica is an open-source question-answering engine designed to offer a self-hosted alternative to Perplexity, with a high degree of flexibility.
- Perplexica supports both cloud and local models, and can call models via the APIs of OpenAI, Anthropic, and Grok, or call open-source LLMs through Ollama.
- Deploying Perplexica requires Docker, and it can be installed with the
docker compose upcommand. - Perplexica’s UI is similar to Perplexity’s, and it supports a Copilot feature that generates multiple search queries based on your prompt to improve results.
- Perplexica also supports cloud deployment, and users can one-click deploy it on platforms like RepoCloud to create a personal question-answering engine.
Full Content
I’ve always wanted to deploy a question-answering engine locally.
In the AI work system I have in mind, a question-answering engine is the foundation. But Perplexity, which is currently the best at this, is pretty picky about the network environment. It’s really annoying when you want to use it and it suddenly stops working.
So a lot of the time, it’s not that I don’t want to pay for SaaS—it’s that these objective conditions force me to take the local deployment route.
Fortunately, there are quite a few projects in this category. I previously introduced one called LLocalSearch. After messing around with a bunch of them, the one I’m most satisfied with right now is Perplexica.
You can tell from the name that this product is basically a Perplexity clone. Put them side by side and the UIs are almost identical.
The main reason I’m satisfied with it is that it’s highly flexible.
On the model side, you can go cloud-based and call the relevant models through the APIs of OpenAI, Anthropic, or Grok. You can also go local and call open-source LLMs through Ollama.
I deleted the ones I had installed before and went through the whole process again, so everyone can see it clearly.
First, open Docker, since we’ll need it in a moment. Then, as usual, use git clone to download the project. After that, remove sample from the front of the config file name.
You can configure the LLM settings in config. For example, fill in the OpenAI API Key, or the Ollama address. If you haven’t changed the port, then it’s the default 11434. One thing to note: don’t enter localhost:11434; enter host.docker.internal:11434, because we’re running inside Docker.
It’s fine if you don’t fill that in here. After everything is installed, you can configure it on the settings page inside the app.
Finally, use the docker compose up command, and it will automatically download and install everything it needs. After a few minutes, you can use it through the local page at localhost:3000.
Let’s test it out. First, let’s try GPT-4o. You can see that it gives a result in about four to five seconds, which is pretty good. The answer sources and follow-up questions are just like Perplexity.
If you turn on the Copilot option, the AI will generate a few more queries based on your prompt and search with them together, which improves the overall result.
Next, let’s try the open-source model. The language model uses qwen2, and the embedding model uses nomic. The first startup is a bit slow because it needs to load things. After that, it’s obviously much faster.
As I said earlier, the main reason I like Perplexica is its flexibility. And that flexibility is not limited to models.
In terms of deployment, besides local deployment, it also supports cloud deployment. On the official GitHub page, there’s a one-click deployment button at the bottom.
It should have a partnership with RepoCloud. After you register there, you’ll get 3 dollars in free credit. At that point, all you need to do is search for the project name and find Perplexica; then enter the OpenAI API Key, as well as the username and password; and finally wait about 5 minutes, and the project will be deployed in the cloud.
As you can see, RepoCloud provides a link, and we can use it freely on desktop and mobile. For example, if I open it on my iPad and log in with the username and password I just set, I’ll see the same interface. Once it’s running, the speed is still OK. RepoCloud will auto-scaling based on your usage.
I found that this kind of personal-exclusive feeling is especially great. I strongly recommend giving it a try. Whether you use it yourself or share it with a team, it works.
OK, that’s it for this episode. Next, I plan to take a closer look at Perplexica and the search engine it uses, SearXNG. If I find anything new, I’ll share it in the newtype community. If you haven’t joined yet, hurry up and join. See you next time!