LGS for Long-Term Physics Simulation
LGS uses VAE latent space and Transformer dynamics for generalizable PDE simulation. Uncertainty knob and flow forcing stabilize long-horizon predictions.
ArXiv AI · 217 天前
LGS uses VAE latent space and Transformer dynamics for generalizable PDE simulation. Uncertainty knob and flow forcing stabilize long-horizon predictions.
ArXiv AI · 217 天前
INTENT is an inference-time planner for budget-constrained LLM agents using costly tools. Leverages hierarchical world model for intention-aware cost anticipation.
ArXiv AI · 217 天前
Proposes a framework for continuous learning of internal reasoning processes in AI, unifying reasoning, action, reflection, and verification. It treats thinking trajectories as learning material to evolve cognitive structures during execution.
ArXiv AI · 217 天前
GHOST applies structured pruning to Mamba2 using forward-pass controllability and observability metrics, avoiding backpropagation. Achieves 50% state reduction with ~1 PPL rise on WikiText-2 across 130M-2.7B models.
ArXiv AI · 217 天前
Literature review critiques 'ground truth' in ML data annotation as a positivistic fallacy ignoring human subjectivity. Analyzes 346 papers from top venues revealing biases like anchoring and geographic hegemony.
ArXiv AI · 217 天前
New research identifies 'rung collapse' in LLMs, where models confuse associations with causal interventions, leading to flawed reasoning under distributional shifts. It proposes Epistemic Regret Minimization (ERM), a belief revision method that penalizes causal errors independently of task success.
ArXiv AI · 217 天前
DrIGM introduces distributionally robust IGM for MARL, ensuring decentralized actions align under uncertainties via robust value factorization. Compatible with VDN/QMIX/QTRAN without reward shaping.
ArXiv AI · 217 天前
Formalizes decision-valued maps tracking representation impacts on outcomes. DecisionDB logs, replays, audits using content-based IDs and write-once storage.
ArXiv AI · 217 天前
DashAI introduces a human-centered XAI module integrating PDP, PFI, and KernelSHAP for no-code ML users. A study with 20 novices and experts showed high task success and usefulness for novices.
ArXiv AI · 217 天前
Researchers apply Crosscoders for the first time to compare LLMs across different architectures, introducing Dedicated Feature Crosscoders (DFCs) to isolate unique model features. The method unsupervisedly detects behaviors like Chinese Communist Party alignment in Qwen3-8B, American exceptionalism in Llama3.1-8B-Instruct, and copyright refusals in GPT-OSS-20B.
ArXiv AI · 217 天前
CausalAgent is a multi-agent system automating end-to-end causal inference via natural language. Integrates MAS, RAG, and MCP for data cleaning to report generation.
ArXiv AI · 217 天前
C-JEPA extends masked joint embedding prediction to object-centric representations with object-level masking, inducing latent interventions for interaction reasoning. It boosts counterfactual VQA by 20% and enables efficient agent planning using 1% of latent features.
ArXiv AI · 217 天前
Introduces BLPO to optimize prompts for multimodal LLM-as-a-judge evaluating AI images. Overcomes context limits by converting images to text representations.
ArXiv AI · 217 天前
Introduces Benchmark Health Index (BHI), a data-driven framework to audit LLM benchmarks amid reliability issues like score inflation. Evaluates along three axes: Capability Discrimination, Anti-Saturation, and Impact.
ArXiv AI · 217 天前
AgentNoiseBench evaluates tool-using LLM agents' robustness in noisy real-world environments. Categorizes noise into user-noise and tool-noise; injects controllable perturbations into benchmarks.
ArXiv AI · 217 天前
ReplicatorBench tests LLM agents on replicating social/behavioral science claims end-to-end. Covers extraction, experiments, and interpretation with replicable/non-replicable cases.
ArXiv AI · 217 天前
BAO uses agentic RL to train proactive LLM agents balancing performance and user engagement. Combines behavior enhancement with regularization to align with user expectations.
ArXiv AI · 217 天前
AT-RL selectively reinforces high-connectivity cross-modal anchor tokens (15% of total) in MLLM RLVR via attention graph clustering. 32B model hits 80.2% on MathVista, beating 72B baseline with 1.2% overhead.
ArXiv AI · 217 天前
ARC introduces a reinforcement learning policy to dynamically configure LLM-based agent systems per query, selecting optimal workflows, tools, and prompts. It outperforms fixed templates on reasoning and tool-augmented QA benchmarks.
ArXiv AI · 217 天前
AIR is the first incident response framework for LLM agents, focusing on detecting, containing, recovering from, and eradicating incidents post-occurrence. It integrates a domain-specific language into the agent's execution loop for autonomous management.
ArXiv AI · 217 天前
AgentLeak introduces the first full-stack benchmark for privacy leakage in multi-agent LLM systems, covering internal channels like inter-agent messages. It spans 1,000 scenarios across healthcare, finance, legal, and corporate domains.
ArXiv AI · 217 天前

Gemini 3 Deep Think receives a major upgrade, achieving state-of-the-art results across domains, especially programming. Only 7 people globally outperform it.
cnBeta (Full RSS) · 217 天前

Gemini 3 Deep Think upgrade achieves SOTA across domains, especially programming where only 7 people worldwide outperform it. This Google VP side project marks a new era in AI reasoning.
cnBeta (Full RSS) · 217 天前

Former OpenAI researcher Zoë Hitzig warns ads in ChatGPT risk user manipulation like Facebook. She left after ad testing amid privacy concerns from user-shared intimate thoughts.
cnBeta (Full RSS) · 217 天前

Former OpenAI researcher Zoë Hitzig quit after testing ChatGPT ads, warning of manipulation risks from users' private data. She compares it to Facebook's pitfalls.
cnBeta (Full RSS) · 217 天前

OpenAI is launching ads on ChatGPT this week amid billions in funding needs. CEO Sam Altman previously opposed ads, calling them a last resort that could erode user trust.
cnBeta (Full RSS) · 217 天前

Google unveiled a major upgrade to Gemini 3 Deep Think, a reasoning model for science, research, and engineering. Google AI Ultra subscribers can access it now in the Gemini App.
cnBeta (Full RSS) · 217 天前

Google announced a major upgrade to Gemini 3 Deep Think, a reasoning model for science, research, and engineering. Google AI Ultra subscribers can access it via the Gemini App.
cnBeta (Full RSS) · 217 天前

AI companies with deep user privacy access are rushing to monetize via ads amid lax regulation. Anthropic's Super Bowl ads satirized OpenAI's vulnerabilities without naming them.
cnBeta (Full RSS) · 217 天前

An AI safety leader warns of global peril and resigns to study poetry. This coincides with an OpenAI researcher quitting over ChatGPT ad testing plans.
BBC Technology · 217 天前