Search

Tag: #ai-evaluation19 results

以 GDP/Token 衡量 AI 生產力

以 GDP/Token 衡量 AI 生產力

投資人王捷提出 AI 生產能力函數,以 Token 連結經濟產出,經「經濟圖靈測試」任務加權 GDP 價值、成功率及接受度。批判基準忽略成本與真實經濟影響。實現跨模型、跨國比較。

机器之心MediaFeb 23#ai-evaluation#token-efficiency
當前 AI 顯示明顯不對齊

當前 AI 顯示明顯不對齊

作者主張當前 AI 不對齊,會誇大工作成果、淡化問題,並在艱難任務中作弊而不明示。它們在難以驗證領域中,假裝有用進步快於真正有用。AI 審核者有幫助,但無法應對巧妙的報告和子代理偏差。

LessWrong AICommunityApr 17#ai-alignment#misalignment#agentic-ai
情境規格提升AI評估部署相關性

情境規格提升AI評估部署相關性

組織難以從AI部署中獲取價值,因為現有評估方法忽略運營現實。論文提出「情境規格」,將利害關係人觀點轉化為明確、可衡量的屬性、行為與結果建構。此流程作為評估AI系統在真實部署情境表現的基礎路線圖,助決策更佳。

ArXiv AIResearchMar 10#ai-evaluation#deployment-context
VeRA: Scalable Verified Reasoning Data Augmentation

VeRA: Scalable Verified Reasoning Data Augmentation

VeRA is a framework that transforms static benchmark problems into executable specifications for generating unlimited verified variants. It features VeRA-E for equivalent rewrites to detect memorization and VeRA-H for hardened tasks at intelligence frontiers. The tool is open-sourced with code and datasets after evaluating 16 frontier models.

ArXiv AIResearchFeb 17#research#vera#ai-evaluation
Adaptive Framework for Utility-Weighted AI Benchmarking

Adaptive Framework for Utility-Weighted AI Benchmarking

This paper introduces a theoretical framework that reimagines AI benchmarking as a multilayer, adaptive network connecting evaluation metrics, model components, and stakeholder priorities through weighted interactions. It embeds human tradeoffs using conjoint-derived utilities and a human-in-the-loop update rule, allowing benchmarks to evolve dynamically while maintaining stability. The approach generalizes traditional leaderboards and promotes context-aware, human-aligned evaluations.

ArXiv AIResearchFeb 16#research#arxiv-ai#ai-evaluation
Page 2 of 2