ArticleOctober 11, 2025Free to read

Everyone Is Miyazaki Hayao

How to use ChatGPT and Hedra to generate a Ghibli-style animated video.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • I introduce how to use ChatGPT and Hedra to generate Ghibli-style animated videos.
  • The workflow includes: using ChatGPT to convert a photo into a Ghibli-style image, then uploading the image and recorded audio to Hedra to generate the video.
  • This method is relatively low-cost (about $30), and since it does not demand much in terms of resolution or lip sync, it is suitable for quickly generating animated content.
  • More precise control can be achieved by generating start and end frame images, then using another tool (such as Kling) to fill in the middle process.
  • I believe AI is making artistic creation more accessible to everyone, and I encourage everyone to actively try and use AI tools for creation.

Full Content

If you also want to generate this kind of Ghibli animation, then you definitely need to watch this episode all the way through.

That segment you just saw was done by me using two tools:

One is ChatGPT. A few days ago, they just added image generation capabilities to GPT-4o. Now, as long as you’re a Plus member, you can use it for $20 a month.

One is Hedra. They are a video generation platform. Pay $10 to become a member, and you get 1,000 usage credits; if that’s not enough, you can buy more.

So, I spent $30 and became a low-budget version of Hayao Miyazaki. The whole creation process is very simple: find a photo and have ChatGPT turn it into a Ghibli-style image. Record a line of speech on your phone and save it as an MP3. Then upload both the image and the audio to Hedra, wait a bit, and it’s delivered.

Let me show you mine. This is the original photo, pasted into GPT. Then I told it to generate a Ghibli style—just those seven characters. After a while, it was done.

Over on Hedra, upload the image. Upload the recorded MP3 too. I just wrote a single simple sentence for the prompt. Then click send. After a few minutes, you can preview and download it.

The advantage of making this kind of style is that it isn’t very sensitive to resolution. For something like mine, both 720P and 4K are acceptable to everyone. As for lip sync, since it’s animation to begin with, it’s fine as long as it’s roughly right. Unlike realistic videos, where even a slight mismatch in lip sync feels uncomfortable.

What I did is only the simplest approach. If you want precise control, you can generate two images—one first frame and one last frame, meaning the beginning and the end. Then send them to Kling, and it will fill in the process with video. You can make it frame by frame like this, and in the end you can create a complete animated work like this.

Ghibli is just one style. You can also use GPT to generate more, like this style too—it’s pretty cool. It all really depends on your taste. Just like I said the other day:

This wave is AI equalizing access, equalizing the right to artistic creation. The last time was the impact of smartphones on photography. The logic is exactly the same.

So, I don’t think there’s much else to say—just do it. I’ve already explained it to this extent, and if you still feel indifferent or keep finding fault with this and that, then I can only wish you good luck.

OK, that’s it for this episode. If you want to discuss AI, come to our newtype community. See you next time!