Search

Tag: #arxiv-ai7 results

Adaptive Framework for Utility-Weighted AI Benchmarking

Adaptive Framework for Utility-Weighted AI Benchmarking

This paper introduces a theoretical framework that reimagines AI benchmarking as a multilayer, adaptive network connecting evaluation metrics, model components, and stakeholder priorities through weighted interactions. It embeds human tradeoffs using conjoint-derived utilities and a human-in-the-loop update rule, allowing benchmarks to evolve dynamically while maintaining stability. The approach generalizes traditional leaderboards and promotes context-aware, human-aligned evaluations.

ArXiv AIResearchFeb 16#research#arxiv-ai#ai-evaluation
Transformers Collapse to Low-Dim Manifolds

Transformers Collapse to Low-Dim Manifolds

Transformer training on modular arithmetic tasks collapses high-dimensional parameters to 3-4D execution manifolds. This structure explains attention concentration, SGD integrability, and sparse autoencoder limits. Core computation occurs in reduced subspaces amid overparameterization.

ArXiv AIResearchFeb 12#research#arxiv-ai#v1
Synthetic Underspecification for Agents

Synthetic Underspecification for Agents

LHAW generates controllable underspecified long-horizon tasks by removing info across goals, constraints, inputs, context. Validates via agent trials, classifying ambiguity impacts. Releases 285 variants from benchmarks.

ArXiv AIResearchFeb 12#research#arxiv-ai#v1
Quadrupeds Cooperate for Super Jumps

Quadrupeds Cooperate for Super Jumps

Co-jump enables two quadrupeds to synchronize jumps up to 1.5m via MAPPO and curriculum, without communication. Achieves 144% height gain over solo robots using proprioception. Transfers from sim to hardware.

ArXiv AIResearchFeb 12#research#arxiv-ai#v1
Diffusion Models Graph Domain Adaptation

Diffusion Models Graph Domain Adaptation

DiffGDA uses diffusion and SDEs to model continuous structure-semantic evolution from source to target graphs. A domain-aware network guides trajectories to optimal adaptation paths. Outperforms baselines on 14 tasks across 8 datasets.

ArXiv AIResearchFeb 12#research#arxiv-ai#v1
1% Params Beat Full Fine-Tuning

1% Params Beat Full Fine-Tuning

CoLin introduces a 1% parameter low-rank complex adapter for vision foundation models. It resolves convergence issues in composite matrices with tailored loss. Surpasses full fine-tuning and delta-tuning on detection, segmentation, and classification.

ArXiv AIResearchFeb 12#research#arxiv-ai#v1