ArticleOctober 10, 2025Free to read

Sora: Standing on OpenAI’s Shoulders

The GPT-3.5 moment for video generation has arrived. The release of Sora marks the point where video generation technology has reached a “usable” level, with realism far surpassing products released around the same time.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • The release of Sora marks the point where video generation technology has reached a “usable” level, with realism far surpassing products released around the same time.
  • The core reason for Sora’s success is that OpenAI adopted the Transformer architecture and applied it to video generation, breaking video into “Spacetime Patches” as tokens.
  • Sora combines the strengths of the Diffusion Model and Transformer, and is known as a “Diffusion Transformer.”
  • During Sora’s training and usage phases, OpenAI fully leveraged its own models like DALL.E 3 and GPT, creating a powerful overall advantage.
  • Sora’s success shows that AI competition has entered a comprehensive arena, where local advantages are hard to withstand against overall leadership, and data will become the key to the next stage of competition.

The GPT-3.5 moment for video generation has arrived.

This technological progress is just way too fast. A year ago, text-to-video looked like this:

This was the very viral “Will Smith eating spaghetti” from back then. It was basically unwatchable, right?

A year later, OpenAI released Sora, and it achieved results like this:

The entire composition, the characters’ skin tones, the lighting and shadows, and so on, all look quite realistic.

If you use the same prompt in Pika, the comparison makes the gap painfully obvious. The others don’t have much time left.

For video generation, there is a very clear threshold between usable and unusable: realism. By realism, I mean whether it conforms to our common sense and the operating laws of the real world, such as the laws of physics.

Look at Sora’s results. This is the first time video generation has reached a usable level. For example, this drone-view clip would fit perfectly into a vlog without any issue.

But rather than marveling at how incredible Sora is, what we should pay more attention to is how exactly OpenAI managed to do all this.

If you’re a practitioner in China, after understanding it, you may feel a bit desperate: is it really possible for us to catch up with OpenAI?

To understand Sora, we need to go back to June 16, 2016. On that day, OpenAI published an article about generative models. The first few paragraphs are crucial:

One of OpenAI’s core aims is to use algorithms and technology to enable computers to understand our world.

To achieve this goal, generative models are one of the most promising paths.

Why insist on “generation”? Feynman once said a very famous line:

If I can’t create it, I don’t understand it.

In other words, if I can generate extremely realistic video, then I must understand the real world well enough.

Look at the title of OpenAI’s latest article:

Video generation models as world simulator.

The idea of treating video generation models as world simulators was established many, many years ago.

And if we look more closely at the technology behind Sora, we can see that everything has been built up little by little over so many years; it’s the inheritance of three generations.

The biggest difference between OpenAI and its peers in developing Sora is that they used the Transformer architecture.

This architecture can be trained on massive datasets, and the cost of fine-tuning is lower as well, so it is especially suitable for large-scale training.

Scalability is the prerequisite for everything OpenAI does. What they want is not academic novelty, but to truly simulate the world and change the world.

Before Transformer, it had already achieved great success in natural language processing. OpenAI believed that one key factor was the use of the concept of tokens.

After text is input, it is split into tokens. Each token is transformed into a vector and then sent to the model. In this way, the Transformer model can use self-attention mechanisms to process and capture complex relationships between tokens, making unified large-scale training convenient.

So when text is replaced by video, tokens become patches.

OpenAI first compresses the video, otherwise the computational load would be too much to handle; then it slices the compressed video into Spacetime Patches.

These patches serve as tokens in the Transformer model, allowing training to proceed just as before.

Sora still belongs to the Diffusion Model category. It takes in low-precision, noisy patches and is trained to predict the original, high-definition patches.

OpenAI calls Sora a Diffusion Transformer because it combines the strengths of both, and that is the technical foundation of Sora’s success.

But that’s not all. Sora is basically a “rich second-generation kid,” with far more resources invested in it than its peers.

During training, video materials need to be paired with text descriptions so the model knows what they are. To improve training quality, OpenAI used its own DALL.E 3 to generate high-quality text descriptions for the video materials.

During usage, the model’s output depends on how precise the user’s prompt is. But you can’t expect users to express themselves clearly in a way that is easy for the model to understand. So OpenAI used its own GPT to expand the user’s prompt in greater detail, and then passed it on to Sora for processing.

So when you put all the factors behind Sora’s success together, you realize this is not at all a case of someone suddenly pulling out a big move:

Text-to-text and text-to-video are supposed to be two different technical paths, right? Yet OpenAI successfully merged them into one.

This shows that in this competition, there are no local battlefields, only a comprehensive arena. Don’t think you can establish a local advantage in some niche and keep the giants out. Pretty desperate, isn’t it?

In the training phase, DALL.E 3 helps with special tutoring; in the usage phase, GPT lends a hand.

Which company’s model gets treatment like that? Pretty desperate, isn’t it?

Large model R&D is moonshot-level difficult. What matters is not talent density, but genius density. Those geniuses have been carrying the grand goal of “letting computers understand the world” and started acting many years in advance. Once they get ahead, it becomes overall leadership.

This is the OpenAI we have to face today.

It will probably take more than half a year before Sora officially goes on sale. For other companies, the question is whether they can replicate this architecture within these eight or nine months. And, very importantly, whether they can find large-scale, high-quality video training data.

Over the past year, everyone has been competing on computing power and algorithms. I feel that the stage where everyone competes on data is about to arrive.