Gemini 3.8 Flash lands on Antigravity with faster coding gains

Antigravity has just rolled out Gemini 3.8 Flash, which Varun Mohan calls a substantial upgrade for agentic coding and knowledge work. Benchmarks show big gains over 3.7 on DeepSWE and mixed results elsewhere. Intro pricing runs through Dec. 31, 2026.

gemini cover

TL;DR

  • Availability: Gemini 3.8 Flash now available to everyone on Antigravity
  • Coding evals: DeepSWE v1.1 73.7% vs 3.7 Flash 65.3%; Terminal-bench 2.1 89.4% vs Opus 5 89.1%
  • Benchmark variability: Terminal-bench 4.0 19.1%, behind Opus 5 51.8% and GPT-5.6 Sol 37.3%
  • General eval gains vs 3.7 Flash: GDPVal-AA v2 1545 Elo, OSWorld-2.0 59.0%, LABBench2 86.2%
  • Not top across set: Opus 5 leads GDPVal-AA v2 (1824 Elo) and OSWorld-2.0 (75.4%); GPT-5.6 Sol leads GDP.PDF (40.0%)
  • Pricing: Intro $0.75/M input, $3.75/M output; regular $1.50/$7.50 from Jan 1, 2027 (intro ends Dec 31, 2026)

Gemini 3.8 Flash is now available to everyone on Antigravity, according to Varun Mohan’s announcement on X. Mohan describes the model as a “substantial improvement” over Gemini 3.7 Flash for agentic coding and general knowledge work, while noting that it had already been used internally.

The accompanying evaluation data places Gemini 3.8 Flash ahead of its predecessor across most listed tests, although the results vary considerably by task. On DeepSWE v1.1, which covers long-horizon software engineering, Gemini 3.8 Flash scores 73.7%, compared with 65.3% for Gemini 3.7 Flash. Its Terminal-bench 2.1 score is 89.4%, narrowly above Claude Opus 5 at 89.1%.

That lead does not carry across every coding benchmark. On Terminal-bench 4.0, Gemini 3.8 Flash records 19.1%, while Claude Opus 5 reaches 51.8% and GPT-5.6 Sol scores 37.3%. The figures suggest that performance depends heavily on the evaluation and workflow involved rather than reflecting a consistent ranking across agentic coding tasks.

Gemini 3.8 Flash also posts higher scores than Gemini 3.7 Flash on several general-purpose evaluations:

BenchmarkGemini 3.8 FlashGemini 3.7 Flash
GDPVal-AA v21545 Elo1482 Elo
Vals Finance Agent v261.4%59.0%
Harvey’s Legal Agent Benchmark10.0%8.8%
CharXiv Reasoning86.2%84.5%
HLE-Verified54.9%53.6%
OSWorld-2.059.0%50.6%
LABBench286.2%82.1%

The model does not lead the full comparison set in all of those categories. Claude Opus 5 scores 1824 Elo on GDPVal-AA v2, 75.4% on OSWorld-2.0 and 90.1% on the Human Solvable version of BioMysteryBench. GPT-5.6 Sol reaches 40.0% on GDP.PDF, compared with 35.0% for Gemini 3.8 Flash.

Antigravity lists Gemini 3.8 Flash and Gemini 3.7 Flash alongside Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol and GPT-5.6 Terra. Both Gemini Flash models have introductory pricing of $0.75 per 1 million input tokens and $3.75 per 1 million output tokens, compared with $1.50 and $7.50 at the regular rates. The introductory prices expire on December 31, 2026; the regular rates begin January 1, 2027.

For comparison, the listed input prices for Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol and GPT-5.6 Terra are $5, $2, $4 and $2 per 1 million tokens, respectively. Their output prices are $25, $10, $20 and $12.

The benchmark methodology is referenced at Google DeepMind’s evaluation methodology page. The available results support Mohan’s claim of improvement over Gemini 3.7 Flash, but they also show that Gemini 3.8 Flash’s standing changes substantially across coding, computer-use, document and specialist research evaluations.

Source: Varun Mohan on X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community