Search

Tag: #research297 results

GT-HarmBench: Game Theory AI Safety Benchmark

GT-HarmBench: Game Theory AI Safety Benchmark

GT-HarmBench introduces 2,009 high-stakes multi-agent scenarios using game theory like Prisoner's Dilemma to benchmark AI safety risks. Frontier models select socially beneficial actions only 62% of the time, often leading to harm. The benchmark, code, and analysis are available on GitHub.

ArXiv AIResearchFeb 16#research#gt-harmbench#ai-safety
EST Boosts Temporal KG Forecasting

EST Boosts Temporal KG Forecasting

Entity State Tuning (EST) introduces persistent entity states to TKG forecasters, overcoming stateless methods' long-term dependency issues. It uses a closed-loop design with topology-aware perception and dual-track evolution. Achieves state-of-the-art results across benchmarks with code on GitHub.

BrowseComp-V³ Benchmark for Multimodal Agents

BrowseComp-V³ Benchmark for Multimodal Agents

BrowseComp-V³ is a new benchmark with 300 challenging questions for evaluating multimodal browsing agents on deep multi-hop reasoning across text and visuals. It features subgoal-driven process evaluation and publicly searchable evidence for reproducibility. Experiments reveal state-of-the-art models achieve only 36% accuracy, highlighting integration bottlenecks.

ArXiv AIResearchFeb 16#research#browsecomp#multimodal-ai
Adaptive Framework for Utility-Weighted AI Benchmarking

Adaptive Framework for Utility-Weighted AI Benchmarking

This paper introduces a theoretical framework that reimagines AI benchmarking as a multilayer, adaptive network connecting evaluation metrics, model components, and stakeholder priorities through weighted interactions. It embeds human tradeoffs using conjoint-derived utilities and a human-in-the-loop update rule, allowing benchmarks to evolve dynamically while maintaining stability. The approach generalizes traditional leaderboards and promotes context-aware, human-aligned evaluations.

ArXiv AIResearchFeb 16#research#arxiv-ai#ai-evaluation
UW Open-Sources MoCo Framework

UW Open-Sources MoCo Framework

University of Washington releases MoCo, a Python framework for multi-model collaboration research. It supports 26 algorithms across API, text, logit, and weight levels. Researchers can customize datasets, models, and hardware to build combinatorial AI systems.

机器之心MediaFeb 16#research#moco#initial
Page 4 of 30