SourceStalecollected in 22h

Gemini 4 Pro Scores Leak Ahead of Launch

Read original on 极客公园
#benchmarks#agent-security#model-competition

Unverified Gemini benchmarks arrive alongside a real warning about AI-assisted attacks and excessive agent permissions.

30-Second TL;DR

What Changed

A leaked image allegedly shows Gemini 4 Pro leading on several agentic benchmarks

Why It Matters

If verified, the benchmark results could intensify competition among frontier model providers. The security incident also demonstrates how AI-assisted exploitation and overprivileged enterprise accounts can amplify a single compromise.

What To Do Next

Do not rely on the leaked Gemini scores; instead, benchmark candidate models in your own harness and audit AI agents’ GitHub, Slack, and cloud permissions.

Who should care:Developers & AI Engineers

Key Points

  • •A leaked image allegedly shows Gemini 4 Pro leading on several agentic benchmarks
  • •The reported scores and testing conditions have not been independently verified
  • •Claude-assisted researchers reportedly reached OpenAI accounts and GitHub access through chained vulnerabilities
Key numbers88%95.3%86.8%$2.25

Deep Insight

Background and context from public sources — not the original article. 9 sources cited.

Enhanced Key Takeaways

  • •The leaked checkpoint surfaced on Arena disguised under the alias 'gemini-3.8-flash' shortly after the legitimate Gemini 3.8 Flash release on September 2, 2026, carrying the internal codename 'Argon'.
  • •Reported agentic and system benchmarks include ~88% on DeepSWE v1.1 (surpassing OpenAI's Astra), 95.3% on Terminal-bench 2.1, 86.8% on OSWorld-2.0, and a 2,064 Elo rating on GDPval-AA v2.
  • •Circulating parameter specifications describe an input context window reaching up to 10 million tokens, an expanded 256K output token limit, and architectural support for persistent cross-session memory.
  • •Leaked API rates list the model at $2.25 per million input tokens and $11.25 per million output tokens, undercutting typical flagship pricing tiers.
  • •Google reportedly canceled Gemini 3.5 Pro after previewing it at Google I/O due to marginal generational improvements, redirecting resources to Gemini 4 based on DeepMind's Recursive Self-Improvement (RSI) research.

Competitor Analysis

Gemini 4 Pro (Argon)
Organization
Google DeepMind
DeepSWE v1.1 Score
~88% (Leaked)
Max Context / Output
10M Input / 256K Output
Leaked / Listed Pricing (per 1M tokens)
$2.25 input / $11.25 output
GPT-6 Astra
Organization
OpenAI
DeepSWE v1.1 Score
~86% (Estimated based on leak gap)
Max Context / Output
Undisclosed frontier context
Leaked / Listed Pricing (per 1M tokens)
Premium frontier pricing
Claude Fable 5.1
Organization
Anthropic
DeepSWE v1.1 Score
Trailing Gemini 4 Pro (Leaked)
Max Context / Output
Extended agentic context
Leaked / Listed Pricing (per 1M tokens)
Premium frontier pricing

Technical Deep Dive

  • Architecture & Training Paradigm: Leverages Google DeepMind's breakthroughs in Recursive Self-Improvement (RSI) rather than relying exclusively on standard compute-scaling runs.
  • Context Window & Memory: Features a 10-million-token input context window, a 256,000-token output limit, and integrated cross-session persistent state memory.
  • Agentic & Execution Metrics: Registered ~88% on DeepSWE v1.1, 2,064 Elo on real-world knowledge evaluation GDPval-AA v2, 95.3% on Terminal-bench 2.1, and 86.8% on OSWorld-2.0 for direct OS navigation.
  • Multimodal Generation Capabilities: Community stress tests revealed zero-shot rendering of dynamic Three.js 3D dashboards and automated generation of functional, interactive canvas/sketching web UIs within 14 minutes.

Future ImplicationsAI analysis grounded in cited sources

Frontier API token pricing will face aggressive downward pressure across major cloud providers.
Offering a 10-million-context flagship model at $2.25 per million input tokens severely undercuts competing frontier models like GPT-6 Astra and Claude Fable 5.1.
Autonomous Recursive Self-Improvement (RSI) will supersede traditional human-curated post-training workflows.
DeepMind's decision to bypass Gemini 3.5 Pro indicates that standard RLHF and synthetic fine-tuning plateaued, necessitating autonomous RSI loops for generational performance leaps.

Timeline

2026-02
Google launches Gemini 3.1 Pro as its prior flagship tier update
2026-05
Google I/O previews Gemini 3.5 Pro before its planned launch is scrapped to focus on Gemini 4
2026-09
Google officially rolls out Gemini 3.8 Flash to general availability
2026-09
Unannounced 'Argon' checkpoint appears on Arena under the alias gemini-3.8-flash, leaking Gemini 4 Pro benchmarks

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 极客公园 ↗

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.