HomeInsightsAI Tools
AI Tools · 9 min read

Google just undercut the cheap AI models on price. The price war is now a three-way fight.

On July 21, 2026, Google launched three new Gemini models built around a single competitive idea: undercut everyone else on cost. Gemini 3.6 Flash is reported as cheaper per task than GPT-5.6 Terra Max, Kimi K3, and Qwen 3.7 Max, alongside a faster Flash-Lite and a security-focused Flash Cyber variant. The strategic point is that Google is now directly attacking the cheap open models on the one dimension they were winning. For a small business running automation, a genuine price war between the largest AI companies is about as favourable a development as you could ask for.

For most of the past year, the story in cheap AI has been about open models from Chinese labs undercutting the expensive Western flagships. GLM-5.2 arrived offering near-frontier capability at a fraction of the price, then Kimi K3 went further and actually beat the closed flagships on a major benchmark while costing far less to run. The implicit question hanging over all of it was how the big Western labs would respond, given that their premium pricing depended on premium capability that was no longer exclusive.

Google answered in late July with three new models designed around cost rather than capability headlines. The framing was explicit: faster, leaner, more token-efficient, and cheaper per task than the competition including the Chinese models that had been winning on price. This is a large, well-resourced company deciding to fight on the dimension where it was being beaten rather than retreating up-market, and the resulting three-way competition between Google, the US labs, and the Chinese open models is genuinely consequential for anyone who pays to run AI at volume.

The five-second answer

Google launched three new Gemini models on July 21, 2026, built explicitly to undercut rivals on cost: Gemini 3.6 Flash, reported as 12 percent faster than its predecessor and cheaper per task than GPT-5.6 Terra Max, Kimi K3, and Qwen 3.7 Max; Gemini 3.5 Flash-Lite, the fastest of the family at around 350 output tokens per second; and Gemini 3.5 Flash Cyber, built to automatically find and patch software vulnerabilities. The strategic significance is that Google is now attacking the cheap Chinese open models on price, the one dimension they were winning. For a small business the meaning is simple and good: three major forces are competing to run your automation more cheaply, so keep your setup swappable and let them fight over your business.

What Google launched

The lineup consists of three models aimed at different points on the speed-cost-capability spectrum. Gemini 3.6 Flash is the workhorse of the group, reported as roughly 12 percent faster than the 3.5 Flash it succeeds and, more importantly, using up to 17 percent fewer tokens for coding and knowledge tasks. That token efficiency matters as much as the headline price, because your actual bill is price per token multiplied by tokens consumed, so a model that does the same job in fewer tokens is cheaper even at identical rates.

Gemini 3.5 Flash-Lite is the speed play, described as the fastest model in the family at around 350 output tokens per second, aimed at workloads where latency matters more than depth. And Gemini 3.5 Flash Cyber is the specialised entry, built to power Google DeepMind's CodeMender agent in automatically scanning, finding, and patching software flaws at a fraction of frontier-model cost, which is an interesting signal that security work is becoming a distinct enough category to warrant its own tuned model.

The competitive claim Google made is the part that matters most for a business. Gemini 3.6 Flash is reported as cheaper per task than GPT-5.6 Terra Max, Kimi K3, and Qwen 3.7 Max, which is a direct pricing shot at both the American flagship tier and the Chinese open models that had been the value leaders. Note the framing is cost per task rather than cost per token, which is the more honest comparison because it accounts for how efficiently a model uses tokens to accomplish something.

The strategy behind it

Google's position going into 2026 was awkward in a specific way. It has enormous resources and genuinely capable models, but it has also been repeatedly slower to market than rivals in several product categories, with its top-tier Pro model timeline slipping while competitors shipped. Rather than trying to win the capability headline race it had been losing on timing, the company appears to be betting that price and efficiency are a more defensible battleground, particularly for the high-volume production workloads where cost dominates.

That bet has real logic behind it. Frontier capability is increasingly commoditised, a point made vividly when Kimi K3 became the first open model to beat the closed flagships outright, which we covered in our Kimi K3 explainer. When several models can all do the job well, the question shifts from which is most capable to which is cheapest to run at scale, and Google has structural advantages in that fight: its own chips, its own data centers, and enormous infrastructure efficiency built over decades.

The security-focused Flash Cyber model also hints at a second strategic idea, which is differentiation by specialisation rather than by raw power. A model tuned specifically for finding and patching software vulnerabilities, running at a fraction of frontier cost, is competing on fit-for-purpose rather than on general intelligence. If that pattern spreads, the market may fragment into cheaper specialised models for particular jobs rather than everyone paying frontier prices for general capability they only partly use.

The three-way price war

What makes this moment genuinely interesting is that there are now three distinct competitive forces pushing prices down simultaneously. The Chinese open models, GLM-5.2 and Kimi K3 among them, compete on being both cheap and freely downloadable, which we compared in detail in our piece on Kimi K3 versus GLM-5.2. The American flagship labs compete on top-end capability and ecosystem. And Google is now explicitly competing on cost-efficiency at volume, attacking the open models where they were strongest.

Three-way competition is meaningfully different from two-way competition because it is harder for any participant to hold a comfortable position. When the cheap option and the capable option were distinct, each could defend its own territory. With a large, well-capitalised player deliberately targeting the cheap end while retaining serious capability, the pressure runs in every direction at once, and none of the participants can easily sit still. That dynamic tends to produce faster price movement than a stable duopoly would.

It is worth being clear-eyed that these competitive claims come substantially from vendor announcements and early reporting rather than from settled independent testing, and cost per task comparisons depend heavily on which tasks you measure. Google says its model is cheaper per task than the named rivals; that claim is plausible and consistent with the token-efficiency numbers, but it is a marketing claim until independently verified across varied real workloads. The competitive direction is clear even where the specific rankings are not.

Why this is good for you

For a small business, an intensifying price war among the world's largest AI companies is close to an unalloyed good. You are the customer they are fighting over, and the mechanism by which they compete is making the thing you buy cheaper and more efficient. Every round of this competition lowers the cost of running the automation that actually helps your business, which expands the set of tasks worth automating and improves the returns on everything you have already built.

The token-efficiency dimension is worth appreciating specifically, because it is a form of price reduction that does not show up in headline rates. A model that accomplishes the same task using fewer tokens costs you less even if the per-token price is unchanged, and Google's claimed 17 percent token reduction on coding and knowledge tasks is a real saving on real bills. As models compete on efficiency rather than just on sticker price, the effective cost of automation falls faster than the published rates suggest.

This continues the trend we have tracked repeatedly, most directly in our guide to what an AI stack actually costs: the cost of capable AI keeps falling, reliably, across every wave of new models. What has changed is the intensity, because with three serious competitive forces rather than one clear leader, the downward pressure is stronger and more sustained. A business building automation now is building on ground that keeps getting cheaper underneath it, which is an unusually favourable position.

Which model should you use?

The honest answer is that it depends on your workload and that you should test rather than trust any comparison, including this one. Google's new Flash models are strong candidates for high-volume production automation where cost and speed dominate, particularly if you already run in Google's ecosystem and the integration is straightforward. Kimi K3 remains compelling where you need maximum capability on demanding or code-heavy work. GLM-5.2 remains the value option for high-volume everyday tasks. The right choice is genuinely workload-dependent.

What should not drive the decision is which model won the most recent benchmark or announcement, because that changes constantly and the differences on your particular tasks may be nothing like the differences on a public leaderboard. The models are close enough in capability for most ordinary business automation that cost, speed, integration convenience, and how they perform on your actual work matter far more than headline rankings. Test with a representative sample of your real workload and let the results decide.

The more important point is that this decision should be low-stakes, and if it feels high-stakes that is a signal about your architecture rather than about the models. In a well-built setup, switching the model behind an automation is a small change you can make in an afternoon, which means you can start with whichever is convenient, test alternatives whenever you like, and move as the price war produces new leaders. Portability turns a difficult choice into an experiment you can repeat cheaply.

What to actually do

Build your automation so the model is a swappable component rather than a hardwired foundation, which is the single highest-value thing you can do in a market moving this fast. When switching models is a configuration change rather than a rebuild, every round of the price war benefits you automatically: a cheaper or better model appears, you test it on your workload, and you move if it wins. Without that portability you are locked into whichever model you happened to pick, watching competitors get cheaper without being able to take advantage.

Periodically re-test your model choice rather than setting it once and forgetting, because in a market where the cost leader changes every few months, a choice made a year ago is very unlikely to still be optimal. This does not need to be constant churn, which has its own costs, but a review every few months of whether your current model is still the right one on cost and quality for your actual workload is a small amount of effort that can meaningfully reduce your bill.

And keep the perspective that the model is not the point. The competitive drama between Google, OpenAI, Anthropic, and the Chinese labs is genuinely making your inputs cheaper, but the value in your business comes from what you automate, not from which engine runs it. The best use of a price war is to let it quietly improve the economics of the automation you build while you concentrate on building the right automation, which is exactly what our 49 euro audit is designed to identify.

The bottom line

Google launching three Gemini models explicitly designed to undercut rivals on cost, including the Chinese open models that had been the value leaders, turns the AI cost competition into a genuine three-way fight between the largest players in the industry. The strategic logic is sound: frontier capability is commoditising, so the durable battleground is cost and efficiency at volume, where Google has real structural advantages in chips and infrastructure.

For a small business this is straightforwardly good news, because you are the customer being fought over and the weapon they are using is lower prices. Every round of this competition makes the automation you run cheaper, expands the set of tasks worth automating, and improves the returns on what you have already built. The right response is not to pick a side but to stay flexible: build so the model behind your automation is swappable, re-test your choice every few months as the leader changes, and keep your real attention on what you are automating rather than which company is currently winning. Let them fight over your business, and quietly bank the savings.

Want automation that stays cheap as the model market shifts? The 49 euro audit sets it up

Sources

Quick answers

Common questions.

Want this in your business?

The €49 audit shows you exactly which automations would pay back fastest in your specific operation.

€49 entryFull AI audit + strategy call included

Reserve your auditNo commitment. No contracts. Just clarity.