Search

Tag: #v1120 results

Confounds Limit FM CT Specificity

Confounds Limit FM CT Specificity

Foundation models match task-specific discrimination in abdominal trauma CT but suffer specificity drops from negative-class heterogeneity like solid organ injuries. Task-specific models handle confounds better. Adaptation via labeled training reduces susceptibility.

ArXiv AIResearchFeb 12#research#foundation-models#v1
CLI-Gym Scales CLI Task Generation

CLI-Gym Scales CLI Task Generation

CLI-Gym generates 1,655 CLI tasks via agentic environment inversion from Dockerfiles. It simulates histories to create buggy states and derives tasks with error messages. Fine-tuned LiberCoder boosts Terminal-Bench scores by 21.1%.

ArXiv AIResearchFeb 12#research#cli-gym#v1
C^2ROPE Advances 3D Multimodal Reasoning

C^2ROPE Advances 3D Multimodal Reasoning

C^2ROPE enhances Rotary Position Embedding for 3D Large Multimodal Models by addressing spatial locality loss and long-term attention decay. It introduces spatio-temporal continuous positional embeddings using triplet hybrid indices and Chebyshev Causal Masking. Evaluations show superior performance on 3D scene reasoning and VQA benchmarks.

ArXiv AIResearchFeb 12#research#c2rope#v1
BNRM Prevents Reward Hacking in RLHF

BNRM Prevents Reward Hacking in RLHF

BNRM introduces Bayesian non-negative reward modeling to combat reward hacking in RLHF. It uses sparse latent factors for disentangled, debiased rewards. Scalable amortized VI enables end-to-end training on LLMs.

ArXiv AIResearchFeb 12#research#bnrm#v1
Blockwise Advantages for Multi-Objective RL

Blockwise Advantages for Multi-Objective RL

Introduces Blockwise Advantage Estimation for GRPO in structured generations, assigning per-objective advantages to avoid interference. Uses Outcome-Conditioned Baseline to estimate advantages without nested rollouts. Competitive on math tasks with uncertainty estimation.

ArXiv AIResearchFeb 12#research#grpo#v1
Benchmark Tests TSFMs on Energy Loads

Benchmark Tests TSFMs on Energy Loads

Multi-dimensional zero-shot benchmark evaluates four TSFMs (Chronos, Moirai, TinyTimeMixer) vs. baselines on ERCOT data. Tests context sensitivity, calibration, robustness to shifts like COVID/Winter Storm. Top models hit MASE 0.31; Chronos-2 best calibrated.

ArXiv AIResearchFeb 12#research#tsfm-benchmark#v1
Benchmark for Self-Evolving Coding LLMs

Benchmark for Self-Evolving Coding LLMs

EvoCodeBench evaluates LLM-driven coding systems on self-evolution, efficiency, and human-comparable performance across languages. Tracks dynamics like solving time and improvements over iterations. Enables cross-language robustness analysis.

ArXiv AIResearchFeb 12#research#evocodebench#v1
Page 11 of 12