ArticleOctober 11, 2025Free to read

Every IP Needs an AI Double, and Every Company Needs an AI Customer Service Agent

Tencent Cloud’s Knowledge Engine for large models drives AI doubles and AI customer service by providing fine-tuned knowledge LLMs, flexible knowledge base settings such as semantic chunking, and search augmentation capabilities.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • The popularization of AI doubles and AI customer service agents is an important sign of AI technology moving into real-world deployment and application exploding, and the participation of cloud vendors is accelerating this process.
  • Tencent Cloud’s Knowledge Engine for large models drives AI doubles and AI customer service agents by providing fine-tuned knowledge LLMs, flexible knowledge base settings such as semantic chunking, and search augmentation capabilities.
  • Knowledge base settings support documents and Q&A sets, and emphasize the importance of evaluation and performance tuning.
  • Tencent Cloud Knowledge Engine’s “workflow management” feature can turn complex processes into tasks executable by AI, enabling highly customized solutions.
  • Knowledge bases and workflows are the core capabilities of agents, corresponding to knowledge and experience, respectively.
  • Tencent Cloud Knowledge Engine also provides atomic capabilities such as multi-turn rewriting, Embedding, Rerank, and document parsing, making integration easier for developers.

Every IP needs an AI double, and every company needs an AI customer service agent. You can remember this sentence of mine and come back in six months to look into it historically.

I’m very sure that in this round of AI technology deployment and AI application explosion, one representative example will be the popularization of AI doubles and AI customer service agents. The former corresponds to super individuals, and the latter to super organizations. This process is accelerating because cloud vendors have already joined in. The market landscape is definitely going to change; it will no longer be a situation dominated by model vendors.

Look, I’ve even added an AI double to my own public account. Behind this agent application is Tencent Cloud’s Knowledge Engine for large models.

I remember that when I first started making videos introducing AI a year ago, RAG tools on the market were especially scarce, and I had to combine and tune all kinds of things myself just to achieve some custom requirements. At one point I even wanted to hand-build a system myself.

If you compare it to now, you’ll find that development over this past year has been incredibly fast. RAG as a Service has appeared, and a whole bunch of out-of-the-box products have emerged. Take the agent application I just mentioned as an example:

For the large model, I’m using the “Fine-Tuned Knowledge LLM Advanced Edition,” with “context rewriting” turned on and the memory turn count increased to 10 turns. You can think of this model as one that has been specially trained for RAG. Of course, if you think the context length is still not enough, you can choose another one, such as the 256K long-text version of Hunyuan LLM. That length is definitely enough.

Just looking at this list, you can see why big companies all need to do foundational model R&D. There are so many business scenarios waiting for specific LLMs to be used. If you don’t seize this strategic initiative yourself, you really must have lost your mind.

For knowledge base settings, I chose “documents,” because these are all ready-made video scripts. If you already have human customer service and want to turn it into AI customer service, there will definitely be Q&A, right? In that case, you can choose “Q&A.”

Generally speaking, Q&A-type materials are more helpful for improving retrieval precision. Later on, I’ll also gradually accumulate a batch of AI-related Q&As, adjusted according to my knowledge reserve and my understanding of AI. The goal is to make this AI double as close to my own cognition as possible.

For recall settings, one is the recall count, meaning how many chunks are recalled and passed to the model; the other is retrieval match threshold, meaning only after similarity reaches a certain value will it be included.

As for the chunk size, the user does not need to set it. Tencent Cloud’s Knowledge Engine will decide on its own where to split based on semantics and the meaning of the entire article, so that it won’t brutally cut off the context. I especially like this point. If you’ve used RAG tools before, you know how troublesome it is to decide chunk size.

Finally, I turned on “search augmentation.” In other words, when the model answers, in addition to referring to the knowledge base I provided, it will also call on the capabilities of WeChat Search and Sogou Search, supplementing more information from within the WeChat ecosystem, such as the massive number of public account articles.

The reason I turned on “search augmentation” is mainly because I don’t want an AI double that only parrots back what it has heard. If your need is AI customer service, then you can leave it off, which makes it more controllable and a bit safer.

Once these basic settings are done, don’t rush to go live. Remember to do evaluation.

First import a sample set, then create an evaluation task. The purpose of evaluation is to see how high the model’s answer accuracy can reach. If the accuracy doesn’t meet the standard, either go back and change the settings, or change the materials.

To be honest, I’ve seen too many people who, after setting up RAG, cursed that it had no effect and that AI was talking nonsense. In fact, in the vast majority of cases, it’s because they naively assumed that simply feeding in all the materials would be enough. In the real world, current technology is not that foolproof yet; you still need to do evaluation and tuning.

Not only that, after formal launch, you’ll also encounter situations where users are dissatisfied with the answers. That’s when “performance tuning” comes into play. On this page, we can see all the answers users were dissatisfied with.

The evaluation I just mentioned is only a simulated scenario, while this is a real business scenario. Only by combining the two can this AI double or AI customer service agent be tuned to the best state. Tencent Cloud being able to think of this and productize it is truly a great merit.

Hello everyone, welcome to my channel. Humbly speaking, I’m one of the few bloggers in China who can explain the Why and How of AI clearly. What I offer is more valuable than tutorials. Remember to follow me. If you want to connect with me, come to the Newtype community. More than 600 friends have already paid to join!

Back to today’s topic: Tencent Cloud’s Knowledge Engine for large models.

A lot of people are focusing on AI applications in the consumer market. I’m actually looking more at the enterprise side, for two reasons:

First, the current AI capabilities are still quite far from market expectations, so it’s hard for consumer-facing products to become phenomenon-level products that solve major problems.

Second, enterprise-side AI has clear demand, the track is very clear, and the returns are quite substantial. So there’s a greater chance of good things emerging here.

For someone like me, a solo operator, using enterprise-grade products is like dimensionality reduction warfare. That’s why I’m optimistic about cloud vendors’ products, and also why I recommend them to everyone.

Tencent Cloud’s Knowledge Engine for large models is a PaaS product. What I just introduced are only the basic RAG functions. If you understand the principles, then this part of the operation should be very easy. If you’re fast, ten minutes is enough.

Going one step further, if you want clearer guidance for this agent application, or if you want to teach your SOP to AI, you must try the workflow management feature.

Here’s a typical example: library customer service. When users look for a library, they generally need three kinds of service: either borrowing books, returning books, or consulting related rules. So on this canvas, you can see three paths, corresponding to three services.

At the beginning of the workflow, AI first makes a conditional judgment based on the user’s inquiry and decides which path to enter. I’ll take borrowing books as an example. Throughout the process, AI will actively guide the user to provide the relevant information.

First, it asks what book to borrow and for how long. Since time is involved, many users express it inconsistently, such as two weeks, one month, and so on, so parameter normalization is needed to unify all expressions into days.

Next, AI will call an interface based on the book title and borrowing duration to check whether it can be borrowed.

If it can be borrowed, it follows the upper branch and asks the user to provide an account ID. If it cannot be borrowed, it follows the lower branch and asks the user whether they want to choose another book.

I’ll demonstrate the conversational effect on the debugging page so everyone can get a feel for it.

Any interaction involving a process can be turned into a workflow. For example, many people ask me how to learn AI. If my AI double handles it, I can use the workflow. Based on my own responses and understanding, I can design a series of conditional judgments and various branch paths, and then teach all of it to AI. So everyone must keep an open mind and not think this whole setup can only be used for customer service.

In addition, one agent application can be connected to N workflows. In other words, you can imagine multiple scenarios and create multiple workflows. AI will autonomously judge which workflow it needs to enter based on the conversation content. This is extremely useful, and the playability is incredibly high!

Knowledge base plus workflows are all the capabilities agents currently have. The former corresponds to knowledge, and the latter corresponds to experience. Tencent Cloud’s Knowledge Engine packages all of this together. So users only need to focus their energy on design, debugging, and invocation.

Design and debugging have already been introduced just now. As for invocation, this Knowledge Engine is mainly API-based, after all it is PaaS. If you have relatively strong development capabilities and needs, and only need part of the engine’s capabilities, you can choose “atomic capabilities,” including:

Multi-turn rewriting is actually for situations where the user’s question may be incomplete. The model will restore it completely by combining contextual semantics. This is quite useful.

Embedding and Rerank: one vectorizes the text, and the other reorders the recalled chunks. Both are essential RAG capabilities.

Document parsing is very basic, very important, and very easy for people to overlook. Good parsing is the starting point of all RAG. Tencent Cloud has a strong advantage here. Many well-known AI products on the market are calling their document parsing technology. They can convert various documents into Markdown format. They can also parse tables, images, as well as content elements like headers, footers, titles, and so on. This really helps a lot and saves us a huge amount of time processing documents.

Tencent Cloud’s Knowledge Engine has very detailed documentation for invoking these four “atomic capabilities,” so I won’t demonstrate them here.

This channel started out by introducing RAG. From using local LLMs to deploying RAG engines, over the past year I’ve shared a lot of content in this area. By the end of the year, vendors have finally launched comprehensive out-of-the-box products. After watching the video, remember to go try Tencent Cloud’s Knowledge Engine.

OK, that’s all for this episode. If you want to discuss AI, come to our Newtype community. See you in the next episode!