CausalAgent: Conversational Causal Inference
CausalAgent is a multi-agent system automating end-to-end causal inference via natural language. Integrates MAS, RAG, and MCP for data cleaning to report generation.
ArXiv AI · 214d ago
Every story we have kept, newest first.
Looking for the daily editions? → Past editions
Page 1360 of 1372
CausalAgent is a multi-agent system automating end-to-end causal inference via natural language. Integrates MAS, RAG, and MCP for data cleaning to report generation.
ArXiv AI · 214d ago
C-JEPA extends masked joint embedding prediction to object-centric representations with object-level masking, inducing latent interventions for interaction reasoning. It boosts counterfactual VQA by 20% and enables efficient agent planning using 1% of latent features.
ArXiv AI · 214d ago
Introduces BLPO to optimize prompts for multimodal LLM-as-a-judge evaluating AI images. Overcomes context limits by converting images to text representations.
ArXiv AI · 214d ago
Introduces Benchmark Health Index (BHI), a data-driven framework to audit LLM benchmarks amid reliability issues like score inflation. Evaluates along three axes: Capability Discrimination, Anti-Saturation, and Impact.
ArXiv AI · 214d ago
AgentNoiseBench evaluates tool-using LLM agents' robustness in noisy real-world environments. Categorizes noise into user-noise and tool-noise; injects controllable perturbations into benchmarks.
ArXiv AI · 214d ago
ReplicatorBench tests LLM agents on replicating social/behavioral science claims end-to-end. Covers extraction, experiments, and interpretation with replicable/non-replicable cases.
ArXiv AI · 214d ago
BAO uses agentic RL to train proactive LLM agents balancing performance and user engagement. Combines behavior enhancement with regularization to align with user expectations.
ArXiv AI · 214d ago
AT-RL selectively reinforces high-connectivity cross-modal anchor tokens (15% of total) in MLLM RLVR via attention graph clustering. 32B model hits 80.2% on MathVista, beating 72B baseline with 1.2% overhead.
ArXiv AI · 214d ago
ARC introduces a reinforcement learning policy to dynamically configure LLM-based agent systems per query, selecting optimal workflows, tools, and prompts. It outperforms fixed templates on reasoning and tool-augmented QA benchmarks.
ArXiv AI · 214d ago
AIR is the first incident response framework for LLM agents, focusing on detecting, containing, recovering from, and eradicating incidents post-occurrence. It integrates a domain-specific language into the agent's execution loop for autonomous management.
ArXiv AI · 214d ago
AgentLeak introduces the first full-stack benchmark for privacy leakage in multi-agent LLM systems, covering internal channels like inter-agent messages. It spans 1,000 scenarios across healthcare, finance, legal, and corporate domains.
ArXiv AI · 214d ago

Gemini 3 Deep Think receives a major upgrade, achieving state-of-the-art results across domains, especially programming. Only 7 people globally outperform it.
cnBeta (Full RSS) · 214d ago

Gemini 3 Deep Think upgrade achieves SOTA across domains, especially programming where only 7 people worldwide outperform it. This Google VP side project marks a new era in AI reasoning.
cnBeta (Full RSS) · 214d ago

Former OpenAI researcher Zoë Hitzig warns ads in ChatGPT risk user manipulation like Facebook. She left after ad testing amid privacy concerns from user-shared intimate thoughts.
cnBeta (Full RSS) · 214d ago

Former OpenAI researcher Zoë Hitzig quit after testing ChatGPT ads, warning of manipulation risks from users' private data. She compares it to Facebook's pitfalls.
cnBeta (Full RSS) · 214d ago

OpenAI is launching ads on ChatGPT this week amid billions in funding needs. CEO Sam Altman previously opposed ads, calling them a last resort that could erode user trust.
cnBeta (Full RSS) · 214d ago

Google unveiled a major upgrade to Gemini 3 Deep Think, a reasoning model for science, research, and engineering. Google AI Ultra subscribers can access it now in the Gemini App.
cnBeta (Full RSS) · 214d ago

Google announced a major upgrade to Gemini 3 Deep Think, a reasoning model for science, research, and engineering. Google AI Ultra subscribers can access it via the Gemini App.
cnBeta (Full RSS) · 214d ago

AI companies with deep user privacy access are rushing to monetize via ads amid lax regulation. Anthropic's Super Bowl ads satirized OpenAI's vulnerabilities without naming them.
cnBeta (Full RSS) · 214d ago

An AI safety leader warns of global peril and resigns to study poetry. This coincides with an OpenAI researcher quitting over ChatGPT ad testing plans.
BBC Technology · 214d ago

A prominent AI safety leader resigned, warning the world is in peril, to study poetry. This follows an OpenAI researcher's exit over plans to test ChatGPT ads.
BBC Technology · 214d ago

Samsung claims first to ship HBM4 memory, a day after Micron's announcement. HBM4 provides faster, denser RAM for next-gen AI hardware.
The Register - AI/ML · 214d ago

Discusses if agent frameworks remain necessary as LLMs improve. Argues building approaches evolve but agents are fundamentally systems around models.
LangChain Blog · 214d ago

Explores whether agent frameworks are still necessary as LLMs improve. Notes that optimal agent-building approaches evolve with model performance.
LangChain Blog · 214d ago

MiniMax released an updated M2 large language model for real-world productivity. The cheap AI follows rivals' launches in China's intense AI race.
SCMP Technology · 214d ago

MIT report shows frontier models like OpenAI's GPT rely on more computing power rather than smarter algorithms. This scaling approach drives progress but hikes costs.
ZDNet AI · 214d ago

Xiaomi open-sources its first-generation VLA large model for robotics. Part of morning tech news roundup alongside OpenAI updates.
Ifanr (爱范儿) · 214d ago

Security flaws in an AI coding platform allowed a BBC reporter to be hacked. Vibe-coding tools enable non-coders to build apps using AI.
BBC Technology · 214d ago

Cloudflare optimizes websites for AI agents by converting HTML to Markdown. Shifts from blocking bots to attracting them with faster content.
The Register - AI/ML · 214d ago

UJET CEO claims AI enhances call agents into 'superheroes' without job loss. Improves software to resolve issues without multi-system navigation.
The Register - AI/ML · 214d ago