The AI model that just scored 65% on DeepSWE isn’t the one Google promised.
Google has a new Gemini model, and no, it is not Gemini 3.5 Pro. Gemini 3.7 Flash launched Thursday as The post The AI model that just scored 65% on DeepSWE isn’t the one Google promised. appeared first on The New Stack .
Google has unveiled Gemini 3.7 Flash, a new AI model designed to excel at coding and handling complex workflows. Although released just three weeks after Gemini 3.6 Flash, the company claims this newer version is superior in writing code and managing longer tasks. To entice developers, Google is offering an introductory API price that is half the original cost of Gemini 3.6 Flash, although this rate will double on January 1, 2027.
Gemini 3.7 Flash demonstrates notable improvements, achieving a 43.6% score on FrontierCode 1.1 Main, up from 34.4% for its predecessor. The model's DeepSWE v1.1 score has also jumped from 49% to 65.3%, reflecting its enhanced performance in software engineering tasks. While coding benchmarks indicate significant progress, the real-world reliability of these improvements within a company's environment remains uncertain.
The model offers three thinking_level settings—low, medium, and high—to balance speed and reasoning. Low is ideal for latency-sensitive tasks, medium strikes a balance for coding and agent workflows, and high allows for more extensive reasoning and tool usage when faced with challenging problems. This flexibility comes at a cost, as longer thinking periods increase token consumption. Currently, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens, with rates doubling on January 1, 2027.
Google's pricing strategy mirrors those of competitors like OpenAI, who have recently reduced their API costs to stay competitive. The discount period provides developers ample time to test Gemini 3.7 Flash at scale, but teams must prepare for the price hike upon its implementation. The company emphasizes the need for planning to accommodate the increased costs, especially for applications with large context or extensive reasoning requirements.
Migrating from previous versions, such as Gemini 3.5 Flash or Gemini 3.1 Pro, involves several changes, including the removal of certain parameters and adjustments to the thinking_budget and candidate_count settings. While the migration process is manageable, engineers should thoroughly test their applications before deploying them in production. Even well-tested code may encounter unexpected behavior when integrated with the new model.
Beyond coding, Gemini 3.7 Flash has shown noteworthy gains in other areas. Its GDP.pdf benchmark improved from 22% to 34%, and its AutomationBench score increased from 17% to 30.4%. However, these results do not guarantee flawless performance in a real-world setting, and teams should exercise caution when deploying the model.
Google's announcement of Gemini 3.7 Flash follows a pattern of delayed model releases across AI labs, with Gemini 3.5 Pro still pending its anticipated June debut in favor of Gemini 4. The ongoing delays highlight the challenges faced by AI developers in delivering cutting-edge models within their promised timelines.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.