Stay connected with KayaToday, follow us on Instagram and Facebook for the latest news and reviews delivered straight to you.
Three weeks. That is how long it took Google to go from Gemini 3.6 Flash to Gemini 3.7 Flash, an interval so short that it says more about Google’s current strategy than any benchmark number does. While the company’s flagship Pro model remains conspicuously absent, Google is iterating furiously on its workhorse Flash line, pairing each upgrade with a pricing move designed to pull enterprise developers deeper into its ecosystem before they fully commit elsewhere.
Gemini 3.7 Flash is now available through the Gemini API in Google AI Studio, Android Studio, and the company’s Antigravity environment. Until December 31, 2026, developers pay $0.75 per million input tokens and $3.75 per million output tokens, exactly half the standard rate of $1.50 and $7.50 that kicks in on January 1, 2027. The discount is temporary by design, giving enterprise teams several months to embed the model into production workflows before the economics reset.
What Actually Improved, and Where the Gaps Remain
Google describes 3.7 Flash as its “most intelligent workhorse model yet for coding and agents,” and its own benchmark data supports that claim in several specific areas, though not universally.
On FrontierCode 1.1 Main, which measures production code quality, 3.7 Flash scores 43.6%, up from 34.4% for its predecessor. Google’s comparison table places that above Claude Sonnet 5 at 42.7% and GPT-5.6 Terra at 41.3%. On DeepSWE v1.1, a long-horizon software engineering evaluation, the model reaches 65.3% versus 49.0% for 3.6 Flash, though GPT-5.6 Terra still leads at 69.6%. Web development benchmarks show 3.7 Flash earning an Elo score of 1588 on Code Arena, ahead of Claude Sonnet 5 at 1541 and GPT-5.6 Terra at 1523.
The picture is more mixed elsewhere. GPT-5.6 Terra leads on Terminal-bench 2.1, Terminal-bench 3.0, and OSWorld-2.0. Claude Sonnet 5 tops the Agent’s Last Exam multimodal desktop tasks with a 33.3% pass rate, compared with 26.3% for 3.7 Flash. Google’s own numbers, in other words, do not show a model that has swept the field. They show one that has become substantially more competitive in coding and document-heavy workflows while sitting at a lower price tier than its main rivals.
The enterprise workflow results are arguably the most interesting. On AutomationBench, which measures enterprise workflow automation, 3.7 Flash scores 30.4%, up sharply from 17.0% for 3.6 Flash. Google’s table lists Claude Sonnet 5 at 10.7% and GPT-5.6 Terra at 23.6%. On GDP.PDF, a complex PDF comprehension evaluation, 3.7 Flash reaches 34.0% versus 22.0% for its predecessor, 28.0% for Claude Sonnet 5, and 24.7% for GPT-5.6 Terra. For agents that must interpret long documents, invoke tools, update external systems, and produce outputs for human review, reliability across that entire chain matters more than isolated reasoning scores.
The Pricing Calculation That Enterprise Teams Actually Need to Run
Comparing token prices in isolation is a trap. A model that costs less per token but requires more retries, more human corrections, or more tool calls to complete a task may end up costing more in practice. Google is betting that 3.7 Flash’s claimed improvements in first-pass accuracy and disciplined multi-step execution will make the math work in its favour even before the introductory discount is factored in.
The introductory pricing makes that bet easier to test. At $0.75 per million input tokens, 3.7 Flash is significantly cheaper than Claude Sonnet 5 at $2 and GPT-5.6 Terra at $2, with output tokens at $3.75 versus $10 and $12 respectively. Context caching costs $0.075 per million tokens during the introductory period. For high-volume coding agents or document-processing pipelines, those differences compound quickly across thousands of daily requests.
The relevant metric for any enterprise team evaluating this model is cost per successfully completed task, not cost per million tokens. Google’s combination of lower introductory pricing and claimed reductions in retries and manual oversight could materially shift that number for the right workloads. Whether it does will depend on testing against actual production repositories, prompts, and failure modes rather than published benchmarks. For developers and technology teams in Malaysia and Singapore building on cloud AI infrastructure, that evaluation is worth running now while the introductory pricing creates a low-cost window to gather real data before committing to a deployment architecture.
The Bigger Problem Google Is Not Talking About
Gemini 3.7 Flash arrives against a backdrop that Google has been careful not to dwell on. The company has not released Gemini 3.5 Pro despite describing it as undergoing partner testing as recently as May, and then saying in July it would become broadly available when ready. Thursday’s announcement offered no timetable. The latest released general-purpose Pro model therefore remains Gemini 3.1 Pro, introduced in February. Reuters reported that 3.5 Pro missed its original target after falling short of internal coding goals, even as Google began training what it calls Gemini 4.
The delay coincides with significant leadership changes. Google DeepMind co-founder Demis Hassabis has stepped back from day-to-day control of the DeepMind division to become its chair and Alphabet’s chief scientist. Former DeepMind CTO Koray Kavukcuoglu now runs the unit as a senior vice president reporting directly to CEO Sundar Pichai, consolidating Gemini model development, frontier research, the Gemini app, and developer teams under a more product-focused operator. Several prominent researchers have departed, including Gemini co-lead Noam Shazeer to OpenAI, AlphaFold scientist John Jumper to Anthropic, and Gemini co-lead Oriol Vinyals along with Jeff Dean and others to a new research startup called Discovery Loop.
Reuters attributed the slower releases partly to internal disagreements, constrained compute allocation, and organisational friction. SemiAnalysis has argued that Google is increasingly prioritising its highly profitable cloud infrastructure business, which serves Gemini competitors, over maintaining absolute frontier model leadership. Google has not confirmed that framing and continues to describe 3.5 Pro as delayed rather than cancelled.
The Verge offered a more measured reading: the departures and delays are serious, but Google retains structural advantages through Search, Workspace, Android, Cloud, custom AI chips, and consumer distribution. The Gemini app has surpassed 950 million monthly users, a reach that does not depend on owning the highest-scoring model on any given leaderboard. Artificial Analysis currently places Gemini 3.7 Flash at 56 on its Intelligence Index, up from 52 for 3.6 Flash but well below Claude Opus 5 at 63. Arena’s early human-preference results are more favourable, provisionally ranking 3.7 Flash ninth overall and eighth for web development.
Why the Flash Strategy Is the Real Story
The three-week release cadence between 3.6 and 3.7 Flash points toward something structurally significant. Google appears to have built a development pipeline in which algorithmic improvements can reach production without waiting for a new flagship generation. The company says the techniques behind 3.7 Flash will inform future models, suggesting this is a deliberate approach rather than an improvised response to competitive pressure.
That cadence creates a genuine tension for enterprise teams. Models can improve quickly, but production deployments require benchmarking against specific repositories, prompt libraries, tool schemas, and failure modes before any change is made. A faster release cycle from Google means more frequent evaluation decisions, not fewer.
Gemini 4 will ultimately determine whether the leadership reorganisation has fixed the execution problems that delayed 3.5 Pro. Until then, Google’s clearest competitive move is the Flash line: models that are competitive enough with more expensive systems, cheap enough to run repeatedly inside agents, and improving fast enough to stay relevant. The introductory pricing on 3.7 Flash is an invitation to test that thesis at low cost before January 2027 resets the economics. Whether the model earns its place in production at full price will depend entirely on how reliably it completes real work, not on where it sits in a comparison table that Google itself compiled.
Read More: Capital One’s AI Playbook: Why the Bank Chose Custom Open Models Over Off-the-Shelf Solutions