🦙Freshcollected in 5h

New Benchmarks Probe Deep Software Engineering Ability

PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#software-engineering#benchmarking#reverse-engineering#code-migrationai-coding-benchmarksgpt-6 astraqwen3.8 27bprogram-benchsre-benchcode migration

💡Traditional coding scores miss reverse engineering and migration; these tests reveal where coding agents truly break.

⚡ 30-Second TL;DR

What Changed

Program-Bench requires agents to recreate a complete codebase from a compiled binary and documentation without decompilers or internet access.

Why It Matters

These benchmarks may provide a more realistic view of agentic software engineering by testing reverse engineering, system understanding, and language migration. They could also expose capability gaps that are hidden by saturated or narrow coding benchmarks.

What To Do Next

Add Program-Bench and Code Migration evaluations to your coding-agent test suite to measure binary understanding and cross-language reimplementation, not just code generation.

Who should care:Researchers & Academics

Key Points

  • Program-Bench requires agents to recreate a complete codebase from a compiled binary and documentation without decompilers or internet access.
  • GPT-6 Astra scored 5.5% on Program-Bench, compared with 5.1% for Fable 5.1 and 0% for Qwen3.8 27B.
  • SRE-Bench measured 88% for GPT-6 Astra, versus 55.9% for GPT-5.6 Sol and 12.5% for Claude Opus 5.
  • On Code Migration, GPT-6 Astra reached 67.7%, while Qwen3.8 27B scored 14.2%.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.