ArticleNovember 6, 2025Free to read

Can Domestic Models Use Skills?

If domestic models are paired with the Claude Code framework, how well do Skills work? I tested three: Zhipu GLM-4.6, Minimax M2, and Moonshot K2.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • Test method: use Claude Code Router to configure domestic models (GLM-4.6/M2/K2) through the ccr ui/model commands, pair them with the Super Analyst Skill to analyze ChatGPT Atlas’ chances, and emphasize following the SOP.
  • Model performance: GLM-4.6 had chaotic tool calls and took multiple tries, M2 failed at search and did not follow the process (incorrect information), K2 handled tool calls smoothly and basically complied (only the Python timing was off).
  • Conclusion: K2 is the best (thanks to its peers for making it look good); domestic models as a whole are weak at following instructions/tools; K2 is recommended, but still stick with the world’s best tools (like Claude) for high cost-performance.

Can domestic models use Skills? This is a question I’ve been asked a lot lately.

Quite a few people in the community, for various reasons, can’t use Claude, so they wonder whether they can use Claude Skills by pairing Claude Code with domestic models.

There are indeed quite a few third-party plugins like this on the market right now, such as Claude Code Router, which I’ve been using all along.

Through the ccr ui command, you can open the backend. I always use OpenRouter as the provider, and it has all the models, which is very convenient. You just need to search for the model name on the official website, then copy and paste it over, and you’re done.

Back in the terminal, use the ccr model command to configure the models. The default model, thinking model, long-context model, search model, and even the image-generation model can all be set in detail.

I made two videos before introducing the first Skill I created—Super Analyst. The Claude model works extremely well with this Skill. For comparison, I used this same Skill and the same question to test three domestic models: Zhipu’s GLM-4.6, Minimax’s M2, and Moonshot’s K2. I’ll show you whether domestic models paired with the Claude Code framework can really use Skills.

Hello everyone, welcome back to my channel. Humbly speaking, I’m one of the few creators in China who can explain both the Why and the How of AI clearly. What I provide is far more valuable than tutorials. Remember to follow me. If you want to connect with me, come join our newtype community. This community has been running for 600 days, and more than 1,800 people have paid to join.

If you’re in China, you can join through Knowledge Planet. If you’re overseas, you can join through Substack. My first course, daily newsletter, and exclusive videos are all available in the community.

Back to today’s topic: domestic models running Skills.

Let’s first take a look at Zhipu’s model.

I set the default model, thinking model, and long-context model all to GLM-4.6.

The question is simple: OpenAI recently launched ChatGPT Atlas. Using the Super Analyst Skill, help me analyze the chances of this AI browser.

Considering these models may not have encountered Skills during training, I specifically added one sentence at the end: please strictly follow the requirements of the Skills.

After running it, you can see that Zhipu’s model did in fact call the Skill and seemed to complete the task. But throughout the process, I didn’t see any MCP calls at all. And the final conclusion was the typical AI-trying-to-fool-you tone. So I followed up with: did you use the Prompt House MCP just now?

Sure enough, it admitted on its own that it had not used the tools as required and had just made up a result on its own.

So I gave Zhipu another chance and had it run it again. I also emphasized once more that it must strictly follow the Skill requirements.

This time it seemed normal. But fortunately, this Skill was made by me. I know the entire SOP very well.

In the framework-selection stage of the analysis, there is a Python script used to help the model make its choice.

It looks like Zhipu’s model ran this script. But the sequence was completely wrong:

It reached a conclusion first, then ran the script to verify whether its framework choice was correct.

In other words, it still did not follow the SOP in the Skill and just made up its own process!

You see, it even admitted it itself: the Python script was used at the wrong time. The order of the steps was chaotic.

It wasn’t until the third run that Zhipu’s GLM-4.6 finally completed the SOP according to the Skill requirements.

Honestly, I’m really, really unhappy with this model. It looks like it can call tools, but once it runs, it’s completely not the same thing. It’s infuriating.

No wonder someone in the community left a comment saying that they first made a Skill and then used prompt text to restrict it so it wouldn’t go off track or bluff. That’s just speechless.

Alright, let’s look at the second model: Minimax’s M2.

Same setup, same question. As soon as I started running it, I had a bad feeling: look, under Web Search, it showed that it did 0 searches.

Sure enough, in the final result, M2 said that OpenAI’s AI browser might not even exist, or that there was extremely limited information.

I was truly speechless. I clearly remember that in their press release, they wrote that it had deep search capabilities and even surpassed a bunch of overseas models, reaching first-tier status.

And throughout the process, it also did not use MCP as required.

I tried again. This time it was even more perfunctory. It just gave the result directly.

Only after I emphasized it did it realize it had to follow the Skill requirements.

I don’t know whether this is an instruction-following problem or a tool-calling problem. It seems that once these models run into a somewhat complex requirement, they become completely at a loss.

Forget it, let’s leave it at that. As a user, I don’t need to think for them. Let’s look at the third model: Kimi K2.

Honestly, after being tormented by the first two models, when I saw K2 smoothly calling tools and formulating a search strategy, I felt a sense of relief—finally, a model with normal behavior!

The entire process basically followed the SOP. The only problem was the timing of that Python script. It should have been run after the search was finished, when it was time to decide which framework to use, not at the very beginning.

How should I put it? There’s a saying: all thanks to the peers making you look good. I find this saying especially, especially applicable domestically.

Think back to the performance of those first two models, and then look at Kimi K2’s performance. All I can say is: it’s already very good.

I remember someone once left a comment asking me why I never introduce domestic models. Now you know why.

If you really have no choice and can only use domestic ones, then try Kimi.

I still stick to my principle: for productivity, only use the world’s best. I believe this is definitely the choice with the highest cost-performance.

OK, that’s it for this episode. If you want to learn AI, want to become a super individual, and want to find like-minded people, come join our newtype community. See you next time!