Here is the number that matters: $0.75 per million input tokens. That is what Google is now charging to run Gemini 3.7 Flash, its newest workhorse coding model โ down from $1.50. Output tokens dropped the same way, from $7.50 to $3.75 per million (InfoWorld).
A clean 50% cut, announced August 13, on a model Google is pitching as its most capable cheap-tier engine for writing code and running agents. The catch โ and there is one โ is that the discount is temporary. On January 1, 2027, the price snaps back to $1.50 and $7.50 (TechTimes).
So this isn't a permanent price war so much as a five-month land grab. Google is renting you a smarter model at cost, betting you won't want to leave once the meter resets. And it aimed the whole thing squarely at the two names developers actually pay for: Anthropic's Claude and OpenAI's GPT.
The thesis: the cheap model just became the one you build on, and the fight has moved from who's smartest to who's cheapest per unit of finished work.
๐ง Why This Matters
For two years the mental model was simple. You reached for the expensive frontier model when the task was hard, and the cheap "Flash" or "mini" tier when you needed speed and didn't care much about quality. Coding lived firmly in the expensive lane.
Gemini 3.7 Flash is Google trying to collapse that split. It scored 43.6% on FrontierCode 1.1, up from 34.4% on the previous 3.6 Flash release just three weeks earlier. On DeepSWE v1.1, a software-engineering benchmark, it jumped to 65.3% from 49.0% (InfoWorld). Those are large moves for a point release, and they land in the cheap tier.
The reason a coding shop cares is arithmetic. When you run agents that loop, retry, and read whole codebases, you burn tokens by the billion. Halving the price of a model that's suddenly competent enough to trust changes what you can afford to automate.
"Token cost has been the practical ceiling on scaling AI beyond isolated pilots." โ Amit Chandak, Chief Analytics Officer, Kanerika (InfoWorld)
๐ Deep Dive
Google's headline pitch is that a cheap model now beats expensive rivals on the thing businesses actually want: finishing real, multi-step workflows without a human babysitting each step. On a benchmark called AutomationBench, here is how Google says the numbers shook out (TechTimes):
- AutomationBench (business workflows): Gemini 3.7 Flash 30.4% โ versus GPT-5.6 Terra at 23.6% and Claude Sonnet 5 at 10.7%. Google's framing: it completes workflows roughly 3x as often as Claude Sonnet 5 and about 30% more often than GPT-5.6 Terra.
- FrontierCode 1.1 (coding): 43.6%, up from 34.4% on 3.6 Flash.
- DeepSWE v1.1 (software engineering): 65.3%, up from 49.0%.
- WebDev Arena (Elo): 1588, up from 1538.
- Artificial Analysis Index (composite): 56, edging Claude Sonnet 5 at 55.
- Price: $0.75 in / $3.75 out per million tokens through 2026, then $1.50 / $7.50.
The pattern is a cheap model posting numbers that used to belong a tier or two up, at a price that undercuts the models it's being compared to.
โ ๏ธ The Catch
Read those benchmarks with your eyebrows raised. They are Google's own results, and one of the flashiest โ AutomationBench โ was published by Zapier, a workflow-automation company with an obvious commercial stake in the idea that AI agents can run business workflows (TechTimes). A 3x lead on a vendor-adjacent test is a marketing number until someone independent reproduces it.
"These remain vendor benchmark claims until the new model accumulates sufficient independent production evidence." โ Sanchit Gogia, CEO, Greyhound Research (InfoWorld)
Even Google's own scorecard doesn't show 3.7 Flash beating pricier competitors everywhere (Slashdot). And the price is the loudest asterisk of all: the $0.75 rate is an introductory discount, not the real sticker. Build a pipeline around it now, and your bill doubles the day 2027 begins.
๐ฏ What Happens Next
Watch two things. First, whether independent evaluations back up the AutomationBench gap once developers run 3.7 Flash on their own messy production tasks instead of curated tests. Vendor benchmarks and real repos rarely agree.
Second, the response. Anthropic and OpenAI are already cutting inference prices, as buyers increasingly ask a blunt question: how much useful work does each dollar of tokens actually buy? (Tech Startups) If Gemini 3.7 Flash forces a matching cut on Claude and GPT's cheap tiers, Google wins even if you never switch โ it just made everyone's margins thinner.
๐งฉ Bigger Picture
There's a quieter tell in the timing. Google shipped a Flash upgrade three weeks after the last one, while the flagship Gemini 3.5 Pro has no announced release date, and CEO Sundar Pichai sidestepped Pro questions on recent earnings calls (InfoWorld). The fast-moving action has shifted to the cheap, high-volume tier โ the one that runs billions of agent calls โ while the headline flagship waits.
That reflects where the money now lives. Frontier bragging rights sell keynotes; cheap tokens run the actual factory floor of AI, the endless background loops writing tests, filing tickets, and refactoring code. Whichever company owns the best price-per-completed-task at scale owns the workload, and Google just moved to own it โ for five months, at cost.
The AI race spent two years asking which model is smartest. Google's answer this week: wrong question. Ask which one you can afford to run a million times before lunch.
Sources
- InfoWorld โ Google cuts Gemini 3.7 Flash prices as enterprise AI economics diverge
- TechTimes โ Google Cuts Gemini 3.7 Flash Price in Half, Claims to Top Claude on Business Workflows
- Slashdot โ Gemini 3.7 Flash Targets Coding and Agents With a 50% Price Cut
- VentureBeat โ Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut
- Tech Startups โ Top Tech News Today, August 14, 2026