Original video in Chinese.
Key Takeaway
- The complexity of PDF format (structure, encoding, information loss) makes AI knowledge bases insufficiently accurate when handling PDFs.
- The key to improving AI knowledge base results is first converting PDFs into Markdown and other formats that make it easier for LLMs to extract text.
- Mathpix is a convenient PDF-to-Markdown tool that supports PDF and image uploads, can export multiple formats, and can OCR LaTeX formulas.
- Marker is another open-source PDF-to-Markdown project that supports multiple languages, formula conversion, and image extraction, and can be deployed locally.
- This article emphasizes the importance of raw data processing for RAG results and recommends two PDF-to-Markdown tools.
Full Content
People often complain that AI knowledge bases are not accurate enough and give irrelevant answers. Sometimes I think about it and feel that AI is actually kind of unfairly blamed, because very likely it is not that it lacks capability, but that the documents you gave it at the beginning are problematic, causing errors or incompleteness when it extracts text. How could the chain of retrieval and generation after that possibly be any good?
Take the most common PDF format, for example. It is no trouble for us to read, but extracting text from it is painful for an LLM.
First, PDF structures are very complex. They contain text, images, tables, as well as font and layout information. It is hard for an LLM to sort all of this out, so naturally it is not easy to extract text from it.
Second, different PDFs may use different character encodings, which can lead to text parsing errors.
Third, even if text is successfully extracted, important information such as paragraphs and headings may be lost, causing misunderstandings of the content.
So, to improve the performance of an AI knowledge base, first convert PDFs into formats that make it easy for an LLM to extract text. In this video, I introduce two tools. One is Mathpix, an off-the-shelf product that I recommended in the newtype community. The other is Marker, which I also recommended in the community earlier. Some friends happened to ask how to deploy it specifically, so I’ll explain that in detail in a moment.
Let’s start with Mathpix.
This product has both desktop and mobile versions. I use the web version. It supports uploading PDFs and images. PDFs are usually papers; images are usually handwritten notes or a teacher’s blackboard writing. After importing the material, it will perform recognition, then either save it in the software as a note for syncing across devices, or export it into formats such as Markdown or Word.
For testing, I uploaded a paper of about 8 pages here, and it contains the most common complex PDF formats. In just a few seconds, Mathpix finished processing it. Then, when I chose to export Markdown, I got an md-format file.
After putting it into Obsidian, you can see that the conversion result is pretty good: content that was originally split into two columns has all been arranged into one column; subtitles, paragraph breaks, tables, and so on are all there.
The reason I chose Obsidian is that its notes are originally in md format, and the AI plugin called Copilot has RAG functionality. Now that I have a PDF-to-Markdown tool, I can handle reading papers, digesting them, and taking notes all in one piece of software.
If you are a STEM student or a researcher, you will definitely love Mathpix—one-click OCR can output LaTeX formulas, which is so convenient. If you have a large number of PDF documents that you want to feed to an LLM as reference material, you can also consider subscribing; it’s less than 5 dollars a month.
A few more words: I personally really like the founder’s thinking behind Mathpix. He proposed a concept called Micro-SaaS, which means starting from a small and focused user pain point and providing an extremely specialized product and features. This kind of focus on niche markets is very suitable for today’s AI era.
OK, Mathpix is the most hassle-free solution. Of course, if you don’t want to spend this little bit of money, that’s fine too—then you can deploy Marker locally for conversion.
Marker is a project I found on GitHub, and it is quite popular. It also converts PDFs into Markdown, supports multiple languages, can convert formulas into LaTeX, can extract images as well, and supports GPU and CPU.
Deployment is easy—same old line: if you have hands, you can do it.
Step one, as usual, create an environment and activate it. I don’t need to explain that.
Step two, install PyTorch. Everyone can go to the official website and choose according to their own situation, then download and install it through the specific command. If CUDA is not installed, then go to old Huang’s place and download one first.
Step three, install Marker. Just use pip install.
After these three steps are done, you can start using it.
According to the guidance on GitHub, we need to run it with one line of command. This line of command is divided into four parts:
The first part, that is, the beginning of the command, tells the machine whether you want to convert one document or multiple documents. If it is just one, use marker single.
The second part tells the machine where the document to be converted is stored, that is, the file path.
The third part tells the machine where to store the document after conversion.
The fourth part is some parameter settings. For example, the default batch is 2, which consumes about 3G of VRAM. The higher this value is set, the more VRAM it needs, and the faster the conversion speed becomes.
Once you understand what this command means, using it every time becomes very simple. If your folder never changes, then you really only need to change the file name.
For the demo, I’ll still use that same paper from earlier so we can compare the results.
Run the command, and you can see the progress bar for each step. Take a look here: Marker first performs checks, then finds the reading order, and finally stores the md file in the specified folder.
Besides the main text, the tables in the paper are extracted separately.
I’ll preview the finished result in VS Code. You can see that the result is pretty good.
However, the official documentation also emphasizes that they cannot guarantee 100% success in extracting formulas and tables, because PDF is just too complex and too weird a thing to make a guarantee about. So after the conversion is complete, I still recommend quickly looking it over and checking it.
If you want to convert multiple documents, the idea is the same: use the command to set the storage location and output location, and you can convert all the PDFs in an entire folder. I won’t demonstrate that here; once you try it once, you’ll understand everything.
OK, that’s it for today. Actually, I mentioned this a long time ago in the community: no matter what RAG tool or technique you use, the first step is always to process the raw data before feeding it in, so that the final result can be guaranteed. If you want to discuss further, come to newtype—I’m always there. See you next time!