Search

Tag: #v1120 results

ERM Fixes Causal Rung Collapse in LLMs

ERM Fixes Causal Rung Collapse in LLMs

New research identifies 'rung collapse' in LLMs, where models confuse associations with causal interventions, leading to flawed reasoning under distributional shifts. It proposes Epistemic Regret Minimization (ERM), a belief revision method that penalizes causal errors independently of task success. Experiments across six frontier LLMs show ERM recovers 53-59% of entrenched errors.

ArXiv AIResearchFeb 13#research#llms#v1
BHI Framework Audits LLM Benchmarks

BHI Framework Audits LLM Benchmarks

Introduces Benchmark Health Index (BHI), a data-driven framework to audit LLM benchmarks amid reliability issues like score inflation. Evaluates along three axes: Capability Discrimination, Anti-Saturation, and Impact. Analyzes 106 benchmarks from 91 models in 2025.

ArXiv AIResearchFeb 13#research#bhi#v1
Wavelet Flows Speed Universe Reconstruction

Wavelet Flows Speed Universe Reconstruction

Cosmo3DFlow uses 3D wavelet transform and flow matching for efficient cosmological inference from N-body simulations. Addresses sparsity via spectral compression, enabling 50x faster sampling than diffusion models. Samples initial conditions in seconds at 128^3 resolution.

ArXiv AIResearchFeb 12#research#cosmo3dflow#v1
VulReaD: KG-Guided Vulnerability Reasoning

VulReaD: KG-Guided Vulnerability Reasoning

VulReaD uses a security knowledge graph and teacher LLM for CWE-consistent vulnerability detection beyond binary classification. Student models are fine-tuned with ORPO for taxonomy-aligned reasoning. Boosts F1 scores significantly on real datasets.

ArXiv AIResearchFeb 12#research#vulread#v1
VLM-Enhanced RL for Autonomous Driving

VLM-Enhanced RL for Autonomous Driving

Found-RL integrates foundation models into RL for end-to-end driving via async batch inference to cut latency. Distills VLM guidance using VMR, AWAG; CLIP rewards shaped by conditional alignment. Lightweight policy matches VLM perf at 500 FPS.

ArXiv AIResearchFeb 12#research#found-rl#v1
VESPO Stabilizes Off-Policy LLM Training

VESPO Stabilizes Off-Policy LLM Training

VESPO introduces variational sequence-level soft policy optimization to tackle training instability in RL for LLMs caused by policy staleness and async execution. It derives a closed-form reshaping kernel for importance weights without length normalization. Experiments demonstrate stable training up to 64x staleness on math benchmarks.

ArXiv AIResearchFeb 12#research#vespo#v1
Versor Revolutionizes Geometric Sequences

Versor Revolutionizes Geometric Sequences

Versor uses Conformal Geometric Algebra (CGA) for sequence modeling with SE(3)-equivariance. Outperforms Transformers on N-body dynamics, topology, and benchmarks with fewer parameters. Offers linear complexity and interpretability via rotors.

ArXiv AIResearchFeb 12#research#versor#v1
V-STAR: Value-Guided RecSys Sampling

V-STAR: Value-Guided RecSys Sampling

V-STAR addresses probability-reward mismatch in generative recsys via value-guided decoding and sibling-relative RL. VED efficiently explores high-potential prefixes; Sibling-GRPO focuses on decisive branches. Outperforms baselines in accuracy and diversity.

ArXiv AIResearchFeb 12#research#v-star#v1
Universal Multimodal Immune System Model

Universal Multimodal Immune System Model

EVA is a cross-species, multimodal foundation model harmonizing transcriptomics and histology for immunology. It shows scaling laws and SOTA on 39 tasks from discovery to clinical trials. Open version released for transcriptomics research.

ArXiv AIResearchFeb 12#research#eva#v1
Page 1 of 12