Search

Tag: #multimodal-agents5 results

UILoop Paradigm for GUI Reasoning

UILoop Paradigm for GUI Reasoning

Proposes UI-in-the-Loop (UILoop) paradigm treating GUI reasoning as cyclic Screen-UI-Action process using MLLMs for better UI element understanding. Introduces challenging UI Comprehension task with three metrics and 26K-sample benchmark. Achieves SOTA in UI understanding and GUI reasoning tasks.

Alibaba Launches Qwen3.5 Multimodal VLM

Alibaba Launches Qwen3.5 Multimodal VLM

Alibaba has introduced the open-source Qwen3.5 series designed for native multimodal agents. The inaugural model is a ~400B parameter vision-language model (VLM) featuring reasoning capabilities via a hybrid mixture-of-experts (MoE) and Gated Delta Networks architecture. It excels in understanding and navigating user interfaces, surpassing previous VLMs.

NVIDIA Developer BlogOfficialFeb 27#multimodal-agents#vision-language#moe-architecture