Initializing portfolio

000

Aravind.
All articles
AI3 min read

Google Halved the Price of Gemini Flash, and Nobody Should Be Surprised

Gemini 3.7 Flash landed three weeks after 3.6 at roughly half the price. The benchmark gains are real; the commoditization signal matters more.

AravindChief Technology Officer & Advisor · AI, Cloud & Cybersecurity
Google Halved the Price of Gemini Flash, and Nobody Should Be Surprised

Google Halved the Price of Gemini Flash, and Nobody Should Be Surprised

Google shipped Gemini 3.7 Flash on 14 August 2026, three weeks after 3.6 Flash. It costs $0.75 per million input tokens and $3.75 per million output tokens — roughly half the previous version. The benchmark gains are real, but the pricing is the news.

The numbers

Google positioned it as its most intelligent workhorse model yet for coding and agents. Against 3.6:

  • FrontierCode 1.1 Main: 43.6%, up from 34.4%
  • DeepSWE v1.1: 65.3%, up from 49.0%
  • AutomationBench: 30.4%, up from 17.0%
  • WebDev Arena Elo: 1588, up from 1538
  • GDP.pdf, covering knowledge-intensive domains: 34.0%, up from 22.0%

AutomationBench nearly doubling is the one that should catch the eye of anyone building agents — and all of it arrived three weeks after the last release, at half the price. For comparison, DeepSeek's V4-Pro runs around $3.96 per million output tokens at peak.

Two cautions worth keeping

Sanchit Gogia of Greyhound Research made the necessary point: these remain vendor benchmark claims until the model accumulates independent production evidence. Benchmarks are marketing until someone runs them on your workload.

Amit Chandak of Kanerika was sharper. The more relevant number for production teams, he said, is token efficiency — and the base model layer is commoditizing. That second half is the part worth underlining.

What commoditization at the base layer means for you

If frontier-adjacent capability keeps arriving every three weeks at falling prices, model choice stops being a strategic decision and becomes a procurement one — and a reversible procurement one at that.

Three practical consequences. Stop architecting around a specific model, because anything hard-coded to one vendor's API surface will need rewriting within a year. Measure token efficiency rather than headline price, since a cheaper model that burns more tokens per task is not actually cheaper. And treat your evaluation harness as the durable asset: models rotate, but the ability to tell whether a swap helped or hurt does not.

The cadence signal

There is a second thing buried in the release. Google has given no timeline for its next Pro model, and the pattern is now visible — Flash updates land frequently, Pro moves slowly.

That tracks where the money is. Most production workloads are not frontier-reasoning workloads; they are high-volume, latency-sensitive and cost-sensitive. Optimising the workhorse tier harder than the flagship tier follows actual demand.

Enterprise agent adoption is still described as measured, with governance as the constraint. This release does not change that, and pricing was never going to be what fixed it.

Source: Google cuts Gemini 3.7 Flash prices as enterprise AI economics diverge and Pro cadence slows — InfoWorld

#Google#Gemini#Enterprise AI#LLM#Model Pricing

Comments

Checking you're human…

Keep reading

Get the next essay first

Checking you're human…

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.