Gemini 3.8 Flash Challenges AI Flagships

💡Gemini 3.8 Flash matches flagship coding benchmarks at a fraction of the cost—but real-world agent tests reveal importan
⚡ 30-Second TL;DR
What Changed
Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens, unchanged from Gemini 3.7 Flash.
Why It Matters
The launch raises the price-performance bar for coding and agent models, especially for workloads where inference cost matters. However, teams should evaluate end-to-end agent reliability rather than relying on headline benchmarks, because tool calls, validation loops, and token usage materially affect cost and quality.
What To Do Next
Run a fixed evaluation set through the Gemini 3.8 Flash API, logging tool calls, output tokens, latency, and task-completion quality before migrating coding agents.
Key Points
- •Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens, unchanged from Gemini 3.7 Flash.
- •It scored 73.7% on DeepSWE v1.1, 89.4% on Terminal-Bench 2.1, and 54.9% on HLE-Verified.
- •Higher reasoning effort increases average DeepSWE output from 107,000 to 143,000 tokens and execution steps from 125 to 166.
- •Gemini 3.8 Flash Cyber is available only to trusted defenders through Google's Fairwind program.
- •Developers reported weaker performance on Terminal-Bench 4.0, OSWorld 2.0, and complex 3D or game-generation tasks than official demos suggested.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.



