Search

Few direct matches — filled in with the latest updates.

Tag: #cua-benchmarks1 results

VLMs as Auditors for Computer-Use Agents

VLMs as Auditors for Computer-Use Agents

CUAAudit introduces Vision-Language Models (VLMs) as autonomous auditors for evaluating Computer-Use Agents (CUAs) via observable interactions. A meta-evaluation of five VLMs across three CUA benchmarks on macOS, Windows, and Linux reveals strong accuracy and calibration, but degradation in complex environments and significant inter-model disagreements.