📄較早收集於 14h

GUI-Owl-1.5 稱霸 20 多項 GUI 基準

GUI-Owl-1.5 稱霸 20 多項 GUI 基準
PostLinkedIn
📄閱讀原文: ArXiv AI
#gui-agents#multi-platform#rl-scalinggui-owl-1.5

💡Open-source GUI agent hits SOTA on 20+ benchmarks across desktop/mobile/web (71.6 AndroidWorld)

⚡ 30-Second TL;DR

有什麼變化

SOTA 分數:OSWorld 56.5、AndroidWorld 71.6、WebArena 48.4

為什麼重要

推進跨平台 GUI 自動化,提升即時代理應用。開源賦能開發者微調 SOTA 模型用於自訂案例。

下一步行動

Clone https://github.com/X-PLUG/MobileAgent and test the online cloud-sandbox demo.

誰應關注:Researchers & Academics

關鍵要點

  • SOTA 分數:OSWorld 56.5、AndroidWorld 71.6、WebArena 48.4
  • 指令/思考變體,尺寸 2B/4B/8B/32B/235B
  • 混合數據飛輪,利用模擬/沙盒環境收集數據
  • MRPO RL 演算法處理多平台長視野任務
  • 開源模型及 GitHub 雲沙盒演示

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 5 個來源。

🔑 增強重點摘要

  • GUI-Owl-1.5 achieves state-of-the-art performance among open-source models on over 20 GUI benchmarks, including 56.5 on OSWorld, 71.6 on AndroidWorld, and 48.4 on WebArena[1].
  • The 8B-Thinking variant excels with 52.9 on OSWorld-Verified, surpassing UI-TARS-2 (53.1) at similar scale and general-purpose models like Qwen3-VL-235B-A22B-Think (38.1)[1].
  • On browser benchmarks, GUI-Owl-1.5-8B-Thinking scores 46.7 on WebArena, 40.8 on VisualWebArena, 78.1 on WebVoyager, and 48.6 on Online-Mind2Web, outperforming other open-source models[1].
  • Thinking variants consistently outperform Instruct versions, especially on long-horizon planning tasks like WebVoyager (69.9 to 82.1) and Online-Mind2Web (41.7 to 48.6)[1].
  • GUI-Owl models (e.g., 7B and 32B) show strong prior results on GUI grounding benchmarks like ScreenSpot-Pro and others, with scores up to 99.0 and 96.1 across tasks[2].
📊 競品分析▸ Show
ModelKey BenchmarksScaleNotes
GUI-Owl-1.5-8B-ThinkingOSWorld-Verified: 52.9, WebArena: 46.7, WebVoyager: 78.1, Online-Mind2Web: 48.68BSOTA open-source multi-platform GUI agent[1]
UI-TARS-2OSWorld-Verified: 53.1Comparable to 8BSlightly higher on OSWorld but narrower scope[1]
Qwen3-VL-235B-A22B-ThinkOSWorld-Verified: 38.1235BGeneral-purpose, lags on GUI tasks[1]
GUI-Owl-7BScreenSpot-Pro variants: 64.8-86.4, grounding: up to 99.07BStrong prior grounding performance[2]
UI-Venus-7BScreenSpot-Pro: 74.67BCompetitive but outperformed by GUI-Owl-32B[2]

🛠️ 技術深入

No specific details on model architecture, Hybrid Data Flywheel, or MRPO RL algorithm found in search results. Benchmarks confirm multi-platform (desktop, mobile, browser) evaluation across automation, grounding, tool calling, memory, and knowledge tasks[1].

🔮 前景展望AI analysis grounded in cited sources

GUI-Owl-1.5 advances open-source GUI agents for cloud-edge collaboration, enabling stronger automation on OSWorld, AndroidWorld, and WebArena, potentially accelerating multi-platform agent deployment while competing with proprietary systems[1].

📎 來源 (5)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. arXiv — 2602
  2. arXiv — 2602
  3. localllm.in — Best Local Llms 24gb Vram
  4. infoq.com — Google Translategemma Models
  5. dl.acm.org — 3747588
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。