Original video in Chinese.
Key Takeaway
- Nano Banana’s advantage: Google’s image generation model has extremely strong instruction-following ability, giving you that “say it and it happens” experience; but most users still describe things in a hit-or-miss way, which creates uncertainty.
- Prompt framework: a JSON structure with 9 parts (shot, subject, environment, lighting, camera, color_grade, style, quality, negatives) to improve controllability and efficiency, making it suitable for production use.
- How to fill it in: use AI to extract descriptions from reference images and fill the framework (handle each part separately, then assemble and adjust); it simply avoids mysticism, and can be adapted for video generation as well. I’ll introduce a community-exclusive video on that.
If you’re using Google’s Nano Banana for image generation, you definitely need to watch this video all the way through. I’ll give you a prompt framework. It’s completely structured and modular. Using it will dramatically improve both controllability and efficiency in image generation!
Nano Banana is Google’s image generation model. After its release, it took the world by storm. I rarely use words like “explosive.” But Google’s banana is really “explosive.” Its instruction-following ability is just too good, giving people a sense of accomplishment, like “say it and it happens.”
Unfortunately, when most people give instructions, they’re still in a hit-or-miss state. Some parts are described in great detail, while others are left out. That creates a lot of uncertainty. And in production environments, what we fear most is exactly this kind of uncertainty.
So my prompt framework is meant to solve this problem. For example, if you want to generate an image like this, how should you describe it? This is my prompt. It uses JSON format and contains nine parts.
First, shot. It defines what the overall shot looks like. For example, this is a macro close-up shot with a centered composition.
Second, subject. It defines the main subject in the image. The subject can be a person or an object. For example, this is a luxurious men’s wristwatch. Its case is polished stainless steel. Its dial is white.
If the subject is an object, you can use sub-items like materials and details to describe it. If the subject is a person, you can use sub-items like age, appearance, and pose to describe it.
Third, environment. It defines what the subject’s environment looks like. For example, what kind of stone slab the watch is placed on. What the background looks like.
Fourth, lighting. It defines what the lighting is like. For example, whether it’s natural light or studio lighting. From what angle the light comes in. Whether it’s sharp or soft.
Fifth, camera. It defines the shooting information. For example, what focal length lens is used, what aperture, and from what angle it’s shot.
Sixth, color_grade. It defines the image’s color style. For example, the contrast should be high, and the saturation can be a little lower.
Seventh, style. It defines the style of the image. For example, hyper-realistic CGI rendering.
Eighth, quality. It defines the quality level of the image. For example, we want 8K resolution with perfect material shaders.
Ninth, negatives. It defines what we do not want to appear in the image. For example, we don’t want branded logos to appear, and we don’t want hands either. This is very important. If you don’t specify it, the AI may very well take the initiative and add them in.
You can see that although this prompt also uses natural language descriptions, it adopts a JSON structure and is divided into nine parts for definition. With a framework like this, the AI can clearly understand exactly what the image in your head looks like.
As long as you define it clearly enough, and generate it multiple times, the results will all be broadly consistent. Only things like this can be used in production.
On the other hand, this framework is also a guideline for you. When you’re conceptualizing, don’t just think of one thing after another. Fill in the blanks properly, and think through each part carefully.
So the question is: now that I have the framework, I really don’t know how to express the details inside it. What should I do?
Easy: let AI copy it for you.
First, go find an image. For example, in this picture, everything except the subject is what you want. Then give the image and the prompt framework to AI and let it fill in the blanks for you. After that, you only need to replace the subject part.
Using this method, you can find matching reference images for the first eight parts in the framework separately—for example, the environment you want, the lighting you want, the style you want. Then let AI generate the corresponding descriptions for each one. In the end, all you need to do is assemble them and make a few adjustments, and you’re done.
It’s that simple. There’s really no need to make it so mystical. This is image generation. Video generation is actually similar too. You just need to make some adjustments to this framework. I’ll make a community-exclusive video to introduce it later.
OK, that’s it for this episode. If you want to learn about AI, want to become a super individual, and want to find like-minded people, come join our newtype community. See you next time!