Gemini 4 Pro Scores Leak Ahead of Launch

Unverified Gemini benchmarks arrive alongside a real warning about AI-assisted attacks and excessive agent permissions.
30-Second TL;DR
What Changed
A leaked image allegedly shows Gemini 4 Pro leading on several agentic benchmarks
Why It Matters
If verified, the benchmark results could intensify competition among frontier model providers. The security incident also demonstrates how AI-assisted exploitation and overprivileged enterprise accounts can amplify a single compromise.
What To Do Next
Do not rely on the leaked Gemini scores; instead, benchmark candidate models in your own harness and audit AI agents’ GitHub, Slack, and cloud permissions.
Key Points
- •A leaked image allegedly shows Gemini 4 Pro leading on several agentic benchmarks
- •The reported scores and testing conditions have not been independently verified
- •Claude-assisted researchers reportedly reached OpenAI accounts and GitHub access through chained vulnerabilities
Deep Insight
Background and context from public sources — not the original article. 9 sources cited.
Enhanced Key Takeaways
- •The leaked checkpoint surfaced on Arena disguised under the alias 'gemini-3.8-flash' shortly after the legitimate Gemini 3.8 Flash release on September 2, 2026, carrying the internal codename 'Argon'.
- •Reported agentic and system benchmarks include ~88% on DeepSWE v1.1 (surpassing OpenAI's Astra), 95.3% on Terminal-bench 2.1, 86.8% on OSWorld-2.0, and a 2,064 Elo rating on GDPval-AA v2.
- •Circulating parameter specifications describe an input context window reaching up to 10 million tokens, an expanded 256K output token limit, and architectural support for persistent cross-session memory.
- •Leaked API rates list the model at $2.25 per million input tokens and $11.25 per million output tokens, undercutting typical flagship pricing tiers.
- •Google reportedly canceled Gemini 3.5 Pro after previewing it at Google I/O due to marginal generational improvements, redirecting resources to Gemini 4 based on DeepMind's Recursive Self-Improvement (RSI) research.
Competitor Analysis
- Organization
- Google DeepMind
- DeepSWE v1.1 Score
- ~88% (Leaked)
- Max Context / Output
- 10M Input / 256K Output
- Leaked / Listed Pricing (per 1M tokens)
- $2.25 input / $11.25 output
- Organization
- OpenAI
- DeepSWE v1.1 Score
- ~86% (Estimated based on leak gap)
- Max Context / Output
- Undisclosed frontier context
- Leaked / Listed Pricing (per 1M tokens)
- Premium frontier pricing
- Organization
- Anthropic
- DeepSWE v1.1 Score
- Trailing Gemini 4 Pro (Leaked)
- Max Context / Output
- Extended agentic context
- Leaked / Listed Pricing (per 1M tokens)
- Premium frontier pricing
| Model | Organization | DeepSWE v1.1 Score | Max Context / Output | Leaked / Listed Pricing (per 1M tokens) |
|---|---|---|---|---|
| Gemini 4 Pro (Argon) | Google DeepMind | ~88% (Leaked) | 10M Input / 256K Output | $2.25 input / $11.25 output |
| GPT-6 Astra | OpenAI | ~86% (Estimated based on leak gap) | Undisclosed frontier context | Premium frontier pricing |
| Claude Fable 5.1 | Anthropic | Trailing Gemini 4 Pro (Leaked) | Extended agentic context | Premium frontier pricing |
Technical Deep Dive
- Architecture & Training Paradigm: Leverages Google DeepMind's breakthroughs in Recursive Self-Improvement (RSI) rather than relying exclusively on standard compute-scaling runs.
- Context Window & Memory: Features a 10-million-token input context window, a 256,000-token output limit, and integrated cross-session persistent state memory.
- Agentic & Execution Metrics: Registered ~88% on DeepSWE v1.1, 2,064 Elo on real-world knowledge evaluation GDPval-AA v2, 95.3% on Terminal-bench 2.1, and 86.8% on OSWorld-2.0 for direct OS navigation.
- Multimodal Generation Capabilities: Community stress tests revealed zero-shot rendering of dynamic Three.js 3D dashboards and automated generation of functional, interactive canvas/sketching web UIs within 14 minutes.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2026-02Google launches Gemini 3.1 Pro as its prior flagship tier update
- 2026-05Google I/O previews Gemini 3.5 Pro before its planned launch is scrapped to focus on Gemini 4
- 2026-09Google officially rolls out Gemini 3.8 Flash to general availability
- 2026-09Unannounced 'Argon' checkpoint appears on Arena under the alias gemini-3.8-flash, leaking Gemini 4 Pro benchmarks
Sources (9)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 极客公园 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.

