Independent GPT-OSS-20B Benchmark Reproduction

💡First independent repro of OpenAI gpt-oss-20b agent benchmarks + open harness
⚡ 30-Second TL;DR
What Changed
Reverse-engineered tools from gpt-oss-20b training distribution via prompts
Why It Matters
Validates OpenAI's gpt-oss-20b agent claims independently, boosting trust in published benchmarks. Provides open tools for community agent development and evaluation.
What To Do Next
Clone https://github.com/borislavmavrin/harmonyagent.git and test gpt-oss-20b on SWE-bench.
Key Points
- •Reverse-engineered tools from gpt-oss-20b training distribution via prompts
- •Built harmonyagent GitHub repo for native message encoding
- •Reproduced OpenAI scores: SWE Verified HIGH 60.4% (pub. 60.7%), MEDIUM 53.3% (53.2%)
- •Achieved 91.7% on AIME25 with tools (pub. 90.4%)
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.