Search

Tag: #gui-agents13 results

GUI-Owl-1.5 Tops 20+ GUI Benchmarks

GUI-Owl-1.5 Tops 20+ GUI Benchmarks

GUI-Owl-1.5 introduces multi-size native GUI agent models (2B-235B) supporting desktop, mobile, browser platforms for cloud-edge collaboration. It sets SOTA on 20+ benchmarks like 56.5 on OSWorld, 71.6 on AndroidWorld, and 80.3 on ScreenSpotPro. Open-sourced with innovations in data flywheel, agent reasoning, and multi-platform RL.

ArXiv AIResearchFeb 20#gui-agents#multi-platform#rl-scaling
RiskWebWorld: Realistic GUI Benchmark for E-commerce Risks

RiskWebWorld: Realistic GUI Benchmark for E-commerce Risks

RiskWebWorld is the first realistic interactive benchmark for GUI agents in e-commerce risk management, featuring 1,513 tasks from production pipelines across 8 domains. It includes challenges like uncooperative websites and partial hijackments, with Gymnasium-compliant infrastructure for scalable evaluation and RL. Evaluations show top models at 49.1% success, highlighting scale's importance over zero-shot grounding.

ArXiv AIResearchApr 17#gui-agents#e-commerce#benchmark
Benchmark Humanizes Mobile GUI Agents

Benchmark Humanizes Mobile GUI Agents

Introduces 'Turing Test on Screen' benchmark modeling agent-detection as MinMax optimization to minimize behavioral divergence. Collects high-fidelity mobile touch dynamics dataset, revealing vanilla LMM agents' detectability due to unnatural kinematics. Establishes AHB with metrics and proposes humanization methods achieving high imitability without utility loss.

HyMEM Supercharges GUI Agents

HyMEM Supercharges GUI Agents

HyMEM is a graph-based memory system for GUI agents that combines discrete symbolic nodes with continuous trajectory embeddings, inspired by human memory. It enables multi-hop retrieval, self-evolution through node updates, and dynamic working-memory refreshing. Experiments demonstrate it boosts Qwen2.5-VL-7B by +22.5%, matching or surpassing Gemini 2.5 Pro Vision and GPT-4o.

ArXiv AIResearchMar 12#gui-agents#graph-memory#vlm-agents
ActionEngine: Programmatic GUI Agents via State Machines

ActionEngine: Programmatic GUI Agents via State Machines

ActionEngine is a training-free framework shifting GUI agents from reactive step-by-step LLM calls to programmatic planning using a two-agent architecture. A Crawling Agent builds updatable state-machine memory through offline exploration, while an Execution Agent synthesizes executable Python programs for tasks. It achieves 95% success on WebArena Reddit tasks with one LLM call, reducing costs 11.8x and latency 2x versus baselines.

Ferret-UI Lite: Tiny On-Device GUI Agent

Ferret-UI Lite: Tiny On-Device GUI Agent

Apple presents Ferret-UI Lite, a compact 3B GUI agent for mobile, web, and desktop platforms. It leverages curated real and synthetic GUI data, chain-of-thought reasoning, and visual tool-use to enhance performance in small on-device models. The paper shares key lessons from its development.

Apple Machine LearningOfficialFeb 17#gui-agents#on-device#chain-of-thought
Step-Level Optimization for Efficient GUI Agents

Step-Level Optimization for Efficient GUI Agents

Proposes an event-driven step-level cascade for computer-use agents, defaulting to small policies and escalating to large models only on detected risks. Features Stuck Monitor for progress stalls and Milestone Monitor for semantic drift in long-horizon GUI tasks. Modular framework layers onto existing agents without retraining.

LAMO: Scalable Lightweight GUI Agents

LAMO: Scalable Lightweight GUI Agents

LAMO framework empowers lightweight MLLMs for GUI automation via multi-role orchestration and task scalability. It features role-oriented data synthesis and two-stage training: Perplexity-Weighted Cross-Entropy SFT for knowledge distillation, plus RL for cooperative exploration. LAMO-3B supports monolithic and MAS execution, excelling as a plug-and-play executor with advanced planners.

ArXiv AIResearchApr 17#gui-agents#lightweight-mllms
Page 1 of 2