MobileWorld Tests Autonomous Phone Agents

π‘MobileWorld shows why phone agents fail when tasks require clarification or tools beyond tapping.
β‘ 30-Second TL;DR
What Changed
The benchmark contains 201 tasks across approximately 20 communication, messaging, and productivity apps.
Why It Matters
The benchmark highlights that mobile agents need more than visual click accuracy: they must handle ambiguity, ask questions, and select external tools. Developers building phone-based agents should therefore evaluate interaction policies and tool orchestration separately from pure GUI completion.
What To Do Next
Add separate user-clarification and MCP tool-use suites to your mobile-agent evaluation, measuring success rate, tool-call count, and task completion time.
Key Points
- β’The benchmark contains 201 tasks across approximately 20 communication, messaging, and productivity apps.
- β’User-interaction tasks require the agent to recognize missing information and ask a user for clarification.
- β’MCP tasks allow agents to use tools such as GitHub and arXiv for data gathering and actions beyond GUI tapping.
- β’The best reported Gemini-3-Pro plus UI-Inst-7B combination averaged about 52%, with major failures on the two new task axes.
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.