Search

Tag: #visual-reasoning12 results

Lang2Act:自湧現語言工具鏈提升VLM視覺推理

Lang2Act:自湧現語言工具鏈提升VLM視覺推理

Lang2Act 透過自湧現語言工具鏈提升視覺語言模型 (VLM) 在 VRAG 中的精細視覺感知,避免僵硬外部工具及影像操作造成的資訊損失。它採用兩階段 RL 框架:第一階段建構可重用動作工具箱,第二階段用於推理。性能提升超過 4%;程式碼在 GitHub 上開放。

ArXiv AIResearchFeb 17#visual-reasoning
7B AdaReasoner Outperforms GPT-5 in Visual Puzzles

7B AdaReasoner Outperforms GPT-5 in Visual Puzzles

AdaReasoner, a 7B model, achieves superior performance on visual reasoning tasks like puzzles by dynamically learning tool selection, timing, and usage. It introduces 'Agentic Vision' with iterative think-act-observe loops, outperforming larger models without massive scaling. Open-source code, models, and paper available on arXiv and GitHub.

机器之心MediaFeb 15#research#adareasoner#7b-model
Page 1 of 2