Harbor Standardizes Large-Scale Agent Evaluation

π‘Compare agent performance across 82 curated tasks without paying the cost of a full benchmark suite.
β‘ 30-Second TL;DR
What Changed
Ports more than 80 agentic benchmarks through reviewed adapters with parity experiments.
Why It Matters
The project could make agent evaluations more comparable by reducing integration differences between benchmarks and harnesses. Its curated task set also offers teams a practical way to run difficult evaluations without the cost of the full benchmark suite.
What To Do Next
Run your agent on Harbor-Index with both Terminus-2 and your native harness to measure harness-sensitive performance and failure modes.
Key Points
- β’Ports more than 80 agentic benchmarks through reviewed adapters with parity experiments.
- β’Evaluates 8 models across 54 benchmarks using Terminus-2 and 3 native harnesses.
- β’Harbor-Index contains 82 difficult, diverse tasks from 29 benchmarks.
- β’No evaluated model-harness configuration exceeds a 30% pass rate; GPT-5.5 with Codex scores 28.0%.
- β’Releases adapters, results, analysis, and Harbor-Index as open-source artifacts.
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.

