ArticleOctober 11, 2025Free to read

Gemini’s Comeback

Google Gemini achieved a “comeback” through its image generation and editing capabilities, delivering a brand-new interactive experience rich in both text and images.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • Google Gemini achieved a “comeback” through its image generation and editing capabilities, delivering a brand-new interactive experience rich in both text and images.
  • Gemini’s native multimodal capability is its core selling point, allowing it to understand and process text, audio, images, and video.
  • Gemini has a killer experience in the consumer market, integrating AI versions of PhotoShop and Meitu Xiuxiu.
  • The Gemini experimental model can directly read YouTube links and use its multimodal capabilities to understand video content.
  • The Gemini app has been updated with access to search history and the Deep Research model, improving its practicality.
  • The article predicts that Gemini will lay the foundation for Google’s dominance in the consumer AI market in 2025.

OpenAI must be pretty nervous now. Because Google updated Gemini two days ago and brought out a comeback-level feature. As usual, I’ll show it first, then explain it.

For example, I asked it to help me create an original Doctor Strange design from scratch, starting with line art and finishing with color, showing each step with an image.

Gemini then started from a concept sketch and outline, completed the line art, refined the details, added color, light and shadow, materials and textures, and magical effects.

To get this level on the first generation, while keeping consistency from start to finish, is seriously impressive!

Let’s try another one. This is a photo Musk posted on Twitter. I pasted it into Gemini and asked it not to change the background, only the expression, and make it a smile.

As you can see, it did a pretty good job. The gaze and crow’s feet are all there. That shows Gemini’s understanding and obedience to instructions, as well as its control over local edits, is pretty good.

Even more impressive, I asked it to give me a braised pork tutorial, with images for each step. It generated corresponding images for every step.

This is Gemini’s newly added image generation and editing capability, available in the experimental version of Gemini 2.0 Flash. If you want to try it, you can use AI Studio or access it through the API.

To be honest, compared with the pros, like SD and Flux, Gemini’s images aren’t especially outstanding. But I think what matters more than being professional is that it has found a way into the mass market.

Fusing image generation with text generation has two advantages.

First, the model’s answers are no longer limited to text; they can be richly illustrated with both text and images.

When an image is needed, it generates one directly. Note, it generates it, rather than searching for an image and dropping it in. It’s like I’m talking and drawing at the same time.

This reminds me of Claude’s Artifacts feature that came out last year. I used a comparison at the time: it’s like a college professor pulling a clean blackboard over while lecturing, writing as they speak.

An experience like that is definitely much better than text alone. Right now it’s text plus images; later on, maybe it can generate short videos and integrate them into the answer. That would absolutely be a killer experience in the consumer market.

Second, users don’t need to switch products; everything can be done in one place.

We all inevitably have some image editing needs from time to time. Gemini now feels like it has integrated AI versions of PhotoShop and Meitu Xiuxiu—it’s just so suitable.

As for heavy-duty products like ComfyUI, they’re very powerful, but the barrier to entry is also very high. Those should be used specifically for professional needs; don’t mix them together with mass-market products.

Once this experimental Gemini model came out, I saw a lot of people already thinking about how to use it to make money.

Think about it: since it has strong obedience to human instructions, you can hand it a script and use it to generate storyboards. Then give the storyboards to a visual model to generate video clips from the images, and finally stitch them together into a complete video.

That makes production even more efficient for self-media creators. You see, the strong never waste time talking nonsense. Unlike the people in the comments, who always think this or that isn’t good enough. They use whatever they have, never complain, and focus on making money.

Back to the point. In addition to image generation, this experimental model can also directly read YouTube links. It doesn’t just extract subtitles from videos; it truly uses multimodal capabilities to “understand” them. In the future, Japanese videos or podcast videos can all be processed by Gemini 2.0 Flash.

This is Gemini’s core selling point all along: native multimodal capability. You can see in the paper that text, audio, images, and video are all input together. Then the model chooses whether to output text or images depending on what’s needed.

Gemini is an autoregressive model. Compared with diffusion models, it has better obedience, and it has been optimized for consistency issues, such as using advanced attention mechanisms and multi-scale generation, solving the inherent shortcomings of the architecture. It took such a long time of accumulation to get this comeback today.

I estimate that in a month or two, this experimental model will be available in the Gemini app. In fact, this round of updates also brought some very practical improvements on the app side.

First, it can access search history.

For example, I asked Gemini: I recently searched for a Microsoft project, but I don’t remember what it was. Then it helped me find it in my search history—it turned out to be Microsoft’s markitdown.

Of course, this feature requires the user’s permission. If you don’t want it, you can turn it off at any time.

Second, the Deep Research model has been updated.

Sure enough, just as I thought before, it has been switched from 1.5 to the latest 2.0. That means stronger reasoning, plus Google’s already insanely good search, making Gemini Deep Research even more useful.

All these features are already out in the open. Imagine if they were integrated into the Android system—I believe that’s only a matter of time—then the AI phone would no longer just be a concept.

So I have a bold idea: in 2025, Gemini will establish Google AI’s dominance in the consumer market.

OK, that’s it for this episode. If you want to learn about AI, come to our newtype community. See you next time!