Search

Tag: #research297 results

CLI-Gym Scales CLI Task Generation

CLI-Gym Scales CLI Task Generation

CLI-Gym generates 1,655 CLI tasks via agentic environment inversion from Dockerfiles. It simulates histories to create buggy states and derives tasks with error messages. Fine-tuned LiberCoder boosts Terminal-Bench scores by 21.1%.

ArXiv AIResearchFeb 12#research#cli-gym#v1
C^2ROPE Advances 3D Multimodal Reasoning

C^2ROPE Advances 3D Multimodal Reasoning

C^2ROPE enhances Rotary Position Embedding for 3D Large Multimodal Models by addressing spatial locality loss and long-term attention decay. It introduces spatio-temporal continuous positional embeddings using triplet hybrid indices and Chebyshev Causal Masking. Evaluations show superior performance on 3D scene reasoning and VQA benchmarks.

ArXiv AIResearchFeb 12#research#c2rope#v1
BNRM Prevents Reward Hacking in RLHF

BNRM Prevents Reward Hacking in RLHF

BNRM introduces Bayesian non-negative reward modeling to combat reward hacking in RLHF. It uses sparse latent factors for disentangled, debiased rewards. Scalable amortized VI enables end-to-end training on LLMs.

ArXiv AIResearchFeb 12#research#bnrm#v1
Blockwise Advantages for Multi-Objective RL

Blockwise Advantages for Multi-Objective RL

Introduces Blockwise Advantage Estimation for GRPO in structured generations, assigning per-objective advantages to avoid interference. Uses Outcome-Conditioned Baseline to estimate advantages without nested rollouts. Competitive on math tasks with uncertainty estimation.

ArXiv AIResearchFeb 12#research#grpo#v1
Benchmark Tests TSFMs on Energy Loads

Benchmark Tests TSFMs on Energy Loads

Multi-dimensional zero-shot benchmark evaluates four TSFMs (Chronos, Moirai, TinyTimeMixer) vs. baselines on ERCOT data. Tests context sensitivity, calibration, robustness to shifts like COVID/Winter Storm. Top models hit MASE 0.31; Chronos-2 best calibrated.

ArXiv AIResearchFeb 12#research#tsfm-benchmark#v1
Benchmark for Self-Evolving Coding LLMs

Benchmark for Self-Evolving Coding LLMs

EvoCodeBench evaluates LLM-driven coding systems on self-evolution, efficiency, and human-comparable performance across languages. Tracks dynamics like solving time and improvements over iterations. Enables cross-language robustness analysis.

ArXiv AIResearchFeb 12#research#evocodebench#v1
AugVLA-3D Boosts VLA with Depth Augmentation

AugVLA-3D Boosts VLA with Depth Augmentation

AugVLA-3D integrates depth estimation from RGB inputs via VGGT to enrich 3D features in vision-language-action models. An action assistant module ensures consistency with control tasks. It enhances generalization and robustness in complex 3D robotic environments.

ArXiv AIResearchFeb 12#research#augvla-3d#v1
AudioRouter Boosts LALMs via RL Tool Use

AudioRouter Boosts LALMs via RL Tool Use

AudioRouter applies RL to teach large audio language models (LALMs) when to use external audio tools, improving fine-grained perception without heavy training. It optimizes a lightweight routing policy while freezing the base model. Achieves big gains on benchmarks with 600x less data than traditional methods.

ArXiv AIResearchFeb 12#research#audiorouter#audio-ai
Page 27 of 30