Iris Search Agents Reach New Open-Source Frontier

π‘Learn how SFT-RL and context management pushed an open-source search agent to leading benchmark results.
β‘ 30-Second TL;DR
What Changed
Iris-mini and Iris-pro are trained at the 35B-A3B and 397B-A17B scales.
Why It Matters
Iris suggests that search-agent performance depends heavily on training data design and inference-time context management, not only on base-model scale. If the promised weights and recipe are released, practitioners could reproduce or adapt a strong open-source baseline for multi-hop web research.
What To Do Next
Reproduce Irisβs evaluation setup on BrowseComp or DeepSearchQA, comparing a ReAct agent with and without inference-time context management before changing your model or tool stack.
Key Points
- β’Iris-mini and Iris-pro are trained at the 35B-A3B and 397B-A17B scales.
- β’Training data is reverse-constructed from web hyperlinks into multi-hop entity-graph questions that cannot be solved by string matching.
- β’The SFT-RL climbing process alternates supervised fine-tuning with reinforcement learning against live search.
- β’With context management enabled, the models score up to 88.6 on BrowseComp, 92.9 on DeepSearchQA, and 56.4 on HLE.
- β’The team plans to release model weights and the complete data, training, and evaluation recipe.
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.