Original video in Chinese.
Key Takeaway
- Large model vendors generally focus on context length, but overlook the limitation of output length.
- At present, large models typically output only 2,000 to 3,000 Chinese characters, mainly because they lack long-text training data.
- By adding long-output data to training, Zhipu significantly improved the model’s output length.
- This article calls on vendors to pay attention to and improve model output length to meet everyday needs.
I’ve never really understood why all the large model vendors are competing on context length, but nobody seems to care about output length.
These days, if you launch a new model version without a 128K context window, you’re almost too embarrassed to greet people with it. But output length — that is, how much the model can reply in one go — seems to have stalled. Two or three thousand characters is about the ceiling.
I did a test with ChatGPT and Claude. My prompt was:
Please help me write a fantasy novel themed around “Black Myth: Wukong.” Center the story on Sun Wukong, and tell a fantasy tale in which the Heavenly Court is corrupt and harmful to all three realms, while Sun Wukong and his demon brothers confront the Heavenly Court and save the people. The novel should be no fewer than 10,000 Chinese characters.
ChatGPT’s performance really disappointed me.
It immediately started slacking off, saying that writing 10,000 characters was too much work, and that it could only help me write part of it and provide a rough outline, with the rest left for me to do myself.
Are today’s AIs really becoming like office workers?
When I asked it to keep writing, ChatGPT started phoning it in. It wrote a few chapters in a perfunctory way, then immediately declared the whole story finished.
Honestly, I was ready to curse.
By contrast, Claude was much better. Everyone should subscribe to Claude instead.
Although it couldn’t output all 10,000 characters at once, Claude at least offered a solution: output by chapter, with each chapter around two or three thousand characters, and the user could give feedback at any time.
That’s the attitude AI should have!
I had Claude write a few chapters. I have to say, its writing was pretty good, and it looked quite decent. If you give it specific guidance, writing something for publication definitely wouldn’t be a problem.
These two examples are very representative. For today’s model products, output length is basically around 2,000 characters.
Why is that?
Zhipu explained it in its paper. The core reason is the lack of long-text training data. In the datasets we use to train large models, very little material is longer than 2,000 characters. So if the model has never seen it and hasn’t been trained on it, naturally it can’t produce it.
To solve this problem, the Zhipu team specifically prepared a long-output dataset, with data lengths ranging from 2K to 32K. They combined it with general-purpose data to form a complete dataset, and used it to fine-tune two models that support 128K context windows: GLM-4-9B and Llama-3.1-8B. The effect was immediate.
I tested it on Google Colab, running the two models separately on an A100 GPU. I used the same fantasy novel task from earlier.
GLM-4-9B did relatively well. I pasted what it wrote into Ulysses for everyone to see. It came to 11,000 characters in total, divided into 13 chapters, starting with an introduction to the world-building and ending with the final decisive battle and the defeat of the Heavenly Emperor.
Llama-3.1-8B didn’t quite meet the target, coming in at just over 8,000 characters. Even so, that still far exceeded the average output of two or three thousand characters.
To be honest, when AI actually wrote the novel, I was still pretty shocked and excited — after all, it was the first time I’d seen output that long. Before, the typical situation was that I’d ask AI to help translate a paper or revise a draft, and it would return only half of it before stopping, which was really annoying and inconvenient.
If 32K context length counts as sufficient, then at least a 5,000-character output length is what can meet everyday needs.
Next, I’m going to try using Zhipu’s training set to fine-tune a few more models. I also sincerely hope that domestic vendors stop blindly chasing ultra-long context windows and treating them as a marketing gimmick, the way smartphone makers chase benchmark scores. It’s time to put output length on the agenda.
OK, that’s it for this episode. If you want to find me, come to the newtype community. See you next time!