Search

直接匹配不多,已補上最新動態。

Tag: #social-reasoning4 results

SocialGrid:具身多代理社交推理基準

SocialGrid:具身多代理社交推理基準

SocialGrid 推出受 Among Us 啟發的具身多代理基準,用於評估 LLM 代理在規劃、任務執行與社交推理的能力。頂尖模型如 GPT-OSS-120B 的任務完成率低於 60%,在導航與欺騙偵測上掙扎。它提供規劃預言機以隔離社交技能,並包含故障分析與競爭排行榜。

ArXiv AIResearchApr 21#multi-agent#embodied-ai#benchmark