GUI-Owl-1.5 稱霸 20 多項 GUI 基準
💡Open-source GUI agent hits SOTA on 20+ benchmarks across desktop/mobile/web (71.6 AndroidWorld)
⚡ 30-Second TL;DR
有什麼變化
SOTA 分數:OSWorld 56.5、AndroidWorld 71.6、WebArena 48.4
為什麼重要
推進跨平台 GUI 自動化,提升即時代理應用。開源賦能開發者微調 SOTA 模型用於自訂案例。
下一步行動
Clone https://github.com/X-PLUG/MobileAgent and test the online cloud-sandbox demo.
關鍵要點
- •SOTA 分數:OSWorld 56.5、AndroidWorld 71.6、WebArena 48.4
- •指令/思考變體,尺寸 2B/4B/8B/32B/235B
- •混合數據飛輪,利用模擬/沙盒環境收集數據
- •MRPO RL 演算法處理多平台長視野任務
- •開源模型及 GitHub 雲沙盒演示
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 5 個來源。
🔑 增強重點摘要
- •GUI-Owl-1.5 achieves state-of-the-art performance among open-source models on over 20 GUI benchmarks, including 56.5 on OSWorld, 71.6 on AndroidWorld, and 48.4 on WebArena[1].
- •The 8B-Thinking variant excels with 52.9 on OSWorld-Verified, surpassing UI-TARS-2 (53.1) at similar scale and general-purpose models like Qwen3-VL-235B-A22B-Think (38.1)[1].
- •On browser benchmarks, GUI-Owl-1.5-8B-Thinking scores 46.7 on WebArena, 40.8 on VisualWebArena, 78.1 on WebVoyager, and 48.6 on Online-Mind2Web, outperforming other open-source models[1].
- •Thinking variants consistently outperform Instruct versions, especially on long-horizon planning tasks like WebVoyager (69.9 to 82.1) and Online-Mind2Web (41.7 to 48.6)[1].
- •GUI-Owl models (e.g., 7B and 32B) show strong prior results on GUI grounding benchmarks like ScreenSpot-Pro and others, with scores up to 99.0 and 96.1 across tasks[2].
📊 競品分析▸ Show
| Model | Key Benchmarks | Scale | Notes |
|---|---|---|---|
| GUI-Owl-1.5-8B-Thinking | OSWorld-Verified: 52.9, WebArena: 46.7, WebVoyager: 78.1, Online-Mind2Web: 48.6 | 8B | SOTA open-source multi-platform GUI agent[1] |
| UI-TARS-2 | OSWorld-Verified: 53.1 | Comparable to 8B | Slightly higher on OSWorld but narrower scope[1] |
| Qwen3-VL-235B-A22B-Think | OSWorld-Verified: 38.1 | 235B | General-purpose, lags on GUI tasks[1] |
| GUI-Owl-7B | ScreenSpot-Pro variants: 64.8-86.4, grounding: up to 99.0 | 7B | Strong prior grounding performance[2] |
| UI-Venus-7B | ScreenSpot-Pro: 74.6 | 7B | Competitive but outperformed by GUI-Owl-32B[2] |
🛠️ 技術深入
No specific details on model architecture, Hybrid Data Flywheel, or MRPO RL algorithm found in search results. Benchmarks confirm multi-platform (desktop, mobile, browser) evaluation across automation, grounding, tool calling, memory, and knowledge tasks[1].
🔮 前景展望AI analysis grounded in cited sources
GUI-Owl-1.5 advances open-source GUI agents for cloud-edge collaboration, enabling stronger automation on OSWorld, AndroidWorld, and WebArena, potentially accelerating multi-platform agent deployment while competing with proprietary systems[1].
📎 來源 (5)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI ↗
每週 AI 簡報
每週一封,可隨時退訂。
