CLI-Gym Scales CLI Task Generation
CLI-Gym generates 1,655 CLI tasks via agentic environment inversion from Dockerfiles. It simulates histories to create buggy states and derives tasks with error messages.
ArXiv AI · 217 天前
CLI-Gym generates 1,655 CLI tasks via agentic environment inversion from Dockerfiles. It simulates histories to create buggy states and derives tasks with error messages.
ArXiv AI · 217 天前
C^2ROPE enhances Rotary Position Embedding for 3D Large Multimodal Models by addressing spatial locality loss and long-term attention decay. It introduces spatio-temporal continuous positional embeddings using triplet hybrid indices and Chebyshev Causal Masking.
ArXiv AI · 217 天前
BNRM introduces Bayesian non-negative reward modeling to combat reward hacking in RLHF. It uses sparse latent factors for disentangled, debiased rewards.
ArXiv AI · 217 天前
Introduces Blockwise Advantage Estimation for GRPO in structured generations, assigning per-objective advantages to avoid interference. Uses Outcome-Conditioned Baseline to estimate advantages without nested rollouts.
ArXiv AI · 217 天前
Multi-dimensional zero-shot benchmark evaluates four TSFMs (Chronos, Moirai, TinyTimeMixer) vs. baselines on ERCOT data.
ArXiv AI · 217 天前
EvoCodeBench evaluates LLM-driven coding systems on self-evolution, efficiency, and human-comparable performance across languages. Tracks dynamics like solving time and improvements over iterations.
ArXiv AI · 217 天前
Proposes causal reward shaping from offline data for continuous RL under confounders. Derives tight value bounds via causal Bellman equation for PBRS.
ArXiv AI · 217 天前
Introduces authenticated workflows as a complete trust layer for enterprise agentic AI, protecting prompts, tools, data, and context. Enforces intent and integrity via cryptography and MAPL policy language.
ArXiv AI · 217 天前
AugVLA-3D integrates depth estimation from RGB inputs via VGGT to enrich 3D features in vision-language-action models. An action assistant module ensures consistency with control tasks.
ArXiv AI · 217 天前
AudioRouter applies RL to teach large audio language models (LALMs) when to use external audio tools, improving fine-grained perception without heavy training. It optimizes a lightweight routing policy while freezing the base model.
ArXiv AI · 217 天前
Aletheia is a math research agent that generates, verifies, and revises solutions using advanced Gemini Deep Think. It achieves milestones like fully AI-generated papers, human-AI collaborations, and solving four open Erdos problems.
ArXiv AI · 217 天前
AI-PACE synthesizes literature to propose a framework for integrating AI into medical education across the learning continuum. It identifies key competencies, curricular approaches, and strategies emphasizing longitudinal integration and interdisciplinary collaboration.
ArXiv AI · 217 天前
Frontier AI models excel in advanced math but consistently fail at multi-digit integer addition. Errors primarily stem from operand misalignment or carry failures, explaining most mistakes in top models like Claude, GPT, and Gemini.
ArXiv AI · 217 天前
AgentTrace instruments LLM agents for structured logging across operational, cognitive, and contextual traces. Provides runtime transparency for security and monitoring in high-stakes settings.
ArXiv AI · 217 天前
Proves LLMs possess predictive partial-world models via task-agnostic affordances for intents. Introduces distribution-robust affordances for multi-task efficiency.
ArXiv AI · 217 天前
AD² analyzes vulnerabilities in end-to-end driving agents like Transfuser to physics, EMI, and digital attacks in CARLA. Driving scores drop up to 99% under threats.
ArXiv AI · 217 天前
Lightweight adapters trained on interpretability artifacts enable reliable self-interpretation in frozen LMs. A simple scalar affine adapter outperforms baselines in feature labeling, topic identification, and implicit reasoning decoding.
ArXiv AI · 217 天前
ADAlign tackles graph domain adaptation by adaptively aligning discrepancies via Neural Spectral Discrepancy (NSD). Uses neural characteristic functions and minimax sampling without heuristics.
ArXiv AI · 217 天前
CoLin introduces a 1% parameter low-rank complex adapter for vision foundation models. It resolves convergence issues in composite matrices with tailored loss.
ArXiv AI · 217 天前

The article questions whether Apple's AI-upgraded Siri will launch before CEO Tim Cook retires. It emphasizes that while delays are tolerable, outright failure is unacceptable.
Ifanr (爱范儿) · 217 天前

Samsung Galaxy S26 is slated for reveal by month's end in a tech news roundup. It may introduce the first 2nm processor in smartphones.
Ifanr (爱范儿) · 217 天前
Introduces a robust, 8-parameter model forecasting >99% AI R&D automation by late 2032. Based on conservative compute growth and algorithmic trends, it predicts 1000x-10M x efficiency gains and 300x-3000x research output by 2035.
AI Alignment Forum · 217 天前
Simplified model forecasts 99% AI R&D automation by late 2032 via compute and algo trends. Uses 8 parameters, conservative assumptions like no full automation.
AI Alignment Forum · 217 天前

Reasoning trace length serves as simple confidence estimator in LLMs to combat hallucinations. Performs comparably to verbalized confidence across models, datasets, prompts.
Apple Machine Learning · 217 天前

Apple researchers demonstrate that reasoning trace length serves as a simple, effective confidence estimator in large reasoning models. It performs comparably to verbalized confidence across models, datasets, and prompts, acting complementarily.
Apple Machine Learning · 217 天前
.png)
Together AI introduces Dedicated Container Inference, a production-grade orchestration for custom AI models. It delivers 1.4x–2.6x faster inference speeds.
Together AI Blog · 217 天前
Hugging Face explores OpenEnv for evaluating tool-using AI agents in practical settings. The post details methodologies for real-world testing.
Hugging Face Blog · 217 天前
Hugging Face blog explores OpenEnv for evaluating tool-using AI agents in practical settings. It highlights real-world applications beyond simulated benchmarks.
Hugging Face Blog · 217 天前

Study maps UX design space for LLM-based computer use agents via two-phase research. Phase 1 reviewed systems and interviewed eight UX/AI practitioners to create taxonomy.
Apple Machine Learning · 217 天前
.png)
Together AI launches Dedicated Container Inference for production-grade orchestration of custom AI models. It delivers 1.4x–2.6x faster inference speeds compared to standard methods.
Together AI Blog · 217 天前