ArticleOctober 11, 2025Free to read

Gemini 2.0: The King of Cost-Effectiveness

Gemini 2.0 is currently the most cost-effective large language model. Its Flash-Lite version is extremely cheap, and the Flash version balances performance, price, and speed.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaway

  • Gemini 2.0 is currently the most cost-effective large language model. Its Flash-Lite version is extremely cheap, and the Flash version balances performance, price, and speed.
  • Gemini 2.0 Pro’s context window has been increased to 2 million, making it suitable for complex reasoning and code generation.
  • The Flash Thinking version has chain-of-thought reasoning capabilities and is suitable for logical reasoning and multi-hop question answering.
  • Gemini 2.0 strikes a balance in performance, stability, speed, and price, becoming my main AI application.
  • This article emphasizes that AI will not replace people, but people who use AI will replace people who do not.

Gemini 2.0 is the most cost-effective large language model in the world, period. I know what you’re thinking—it’s even stronger than DeepSeek. Let’s go straight to the prices; overseas bloggers have already made the table.

Gemini 2.0 Flash-Lite is really dirt cheap: input is only 0.075 USD, output 0.3 USD.

Flash, which has a little more functionality than it, is only a bit more expensive: input 0.1 USD, output 0.4 USD.

Now look at DeepSeek: V3 input 0.27, output 1.1; R1 input 0.55, output 2.19.

Google is really pushing hard here. You have to know that Flash is a state-of-the-art model that supports multimodality, native tool use, and a one-million-token context window. Yet it’s not only cheaper than DeepSeek, it also beats GPT-4o mini. Looks like the AI competition this year is going to intensify.

Hello everyone, welcome to my channel. Humbly speaking, I’m one of the few creators in China who can clearly explain the Why and How of AI. What I provide is more valuable than tutorials. Remember to follow. If you want to connect with me, come to the newtype community. More than 800 friends have already joined and paid!

Back to today’s topic: the king of cost-effectiveness—Gemini 2.0.

Gemini 2.0 is the model family Google updated a few days ago, including the Pro and Flash lines.

Pro is easy to understand: it’s Google’s current top-tier model. It has all the features it should have, and it has increased the context window from one million to two million. So the Pro version is very suitable for complex reasoning, code generation, and so on.

Flash, on the other hand, balances performance, price, and speed, and is the main model for daily use. Among them, Flash also has two variants:

Flash-Lite removes a little functionality, such as not supporting image and audio output, not supporting online search, and not being able to execute code, and then pushes the price down to the minimum. So if you need to generate text at scale, Lite is the most suitable version.

As the name suggests, Flash Thinking is the version with chain-of-thought reasoning capability. Just like the familiar DeepSeek-R1, it performs multi-step reasoning before answering. So for some complex tasks, such as those requiring stronger logical reasoning or multi-hop question answering, Flash Thinking is the most suitable.

I said earlier that Gemini 2.0 is the king of cost-effectiveness, but I feel that’s not quite accurate. Because the three words “cost-effectiveness” make it seem like its performance isn’t that great. A more fitting name would be “the king of competition.” Let me show you the results.

Let’s first look at Pro’s capability. My question was:

Why was Nvidia’s CUDA so successful? How deep is its moat really? In the AI era, is it possible for Nvidia’s competitors to catch up or overturn it?

As you can see, although Pro is slower than Flash, it still feels pretty fast. And the answer it gives is very clear logically, with not much unnecessary filler, which I really like.

Now let’s look at Flash Thinking. I asked it a question that has been discussed a lot recently:

Does the success of DeepSeek-R1 mean that Nvidia’s high-compute GPUs and CUDA are no longer needed?

The thinking process of Flash Thinking is in English. First it broke down my question and concluded that it needed to search and research certain keywords, and then it carried out the corresponding search. Just like Pro, its answer is quite clean and refreshing.

For comparison, I asked DeepSeek-R1 the same question. Although the conclusion was basically the same—that it remains irreplaceable, but dependence may decrease—the thinking process was quite different:

Flash Thinking first breaks the problem down, then searches. R1 searches directly first, then looks at what the pages it found are saying. From a methodological perspective, I personally prefer breaking things down first. What do you think?

Gemini is my main AI application. I originally used ChatGPT. But during use, I ran into all kinds of limits, which were really annoying. Just then, Claude 3.5 came out, so I switched to Claude. Later, Claude got massively banned, and all three of my accounts were hit, so I “fled” to Gemini and topped up there too.

With this 2.0 update, I’ve been using it for the past few days and I’m extremely satisfied. No matter which version it is, it achieves a balance of performance, stability, speed, and price. On desktop, using the web version, you have Pro, Flash, and Flash Thinking. On mobile, you can use the official app and choose Pro or Flash.

As long as Google doesn’t pull any tricks, until the next major model update, Gemini will continue to be my daily mainstay.

I know using these overseas products means crossing several barriers. But these barriers actually filter out a lot, a lot of people for you. As long as you spend a little time and a little money to get past them, you gain a huge lead. As I always say:

AI won’t replace you. People who use AI, especially people who use advanced AI, will.

OK, that’s it for this episode. If you want to learn more about AI, come join our newtype community. See you next time!