All Updates

Page 1711 of 1963

March 9, 2026

📄
ArXiv AI173d ago

RoboLayout: Agent-Friendly 3D Scenes

RoboLayout extends LayoutVLM with agent-aware reasoning and reachability constraints for generating navigable 3D indoor scenes from language. It supports diverse agents like robots, humans, and animals via abstract representations. Local refinement optimizes problematic placements efficiently while preserving semantic coherence.

#embodied-ai#3d-generation
📄
ArXiv AI173d ago

RL for Climate-Resilient Transport

Proposes a reinforcement learning (RL) framework for long-term flood adaptation in urban transport systems. Integrates rainfall projections, flood modeling, transport simulation, and impact quantification as an integrated assessment model (IAM). Outperforms traditional optimization in Copenhagen case study by discovering coordinated strategies under climate uncertainty.

#climate-adaptation#flood-modeling#urban-transport
📄
ArXiv AI173d ago

Reasoning Models Fail CoT Control

New CoT-Control suite reveals reasoning models have low controllability over chain-of-thought vs. final outputs, e.g., Claude Sonnet 4.5 at 2.7% CoT control but 61.9% output. Controllability drops with larger models, RL training, compute, and difficulty. Results suggest CoT monitoring remains viable despite evasion attempts.

#chain-of-thought#controllability#model-safety
📄
ArXiv AI173d ago

ProEvolve: Programmable Agent Benchmark Evolution

ProEvolve is a graph-based framework for scalable, controllable evolution of agent environments to test LLM agents' adaptability to real-world changes. It uses a typed relational graph to unify data, tools, and schemas, enabling updates via graph transformations that propagate coherently. The system generates 200 environments and 3,000 task sandboxes for benchmarking.

#agent-benchmarks#graph-framework
📄
ArXiv AI173d ago

New Aggregative Semantics for QBAF

Researchers propose aggregative semantics for Quantitative Bipolar Argumentation Frameworks (QBAF), aggregating attackers and supporters separately unlike modular approaches. This involves a three-stage process: computing global attacker and supporter weights before combining with intrinsic argument weights. The work explores properties, principles, and compares 500 variants for diverse behaviors.

#bipolar-semantics#gradual-aggregation
🗾
ITmedia AI+ (日本)173d ago

Mistral Launches Local Japanese Speech AI Voxtral 2

Mistral AI announced Voxtral Transcribe 2, a speech recognition model that runs locally and supports Japanese. It includes two variants: one for high-accuracy, low-cost batch processing and another for ultra-low latency real-time transcription.

#speech-recognition#on-device#japanese-support
📄
ArXiv AI173d ago

MACRO: Self-Evolving Medical Imaging Agent

MACRO is a self-evolving agent that autonomously discovers reusable composite tools from verified execution trajectories in medical imaging tasks. It grounds tool selection in visual-clinical context via image-feature memory and reinforces reliable use through GRPO-like training. Experiments show superior multi-step accuracy and cross-domain generalization over baselines.

#medical-imaging#agentic-ai#self-improvement
📄
ArXiv AI173d ago

DeepFact: Co-Evolving Benchmarks for LLM Factuality

DeepFact proposes Audit-then-Score (AtS) to create evolving benchmarks for verifying factuality in deep research reports from search-augmented LLM agents. It introduces DeepFact-Bench, a versioned benchmark with auditable rationales, and DeepFact-Eval, a verification agent that outperforms existing tools. Expert accuracy on claims rises from 60.8% to 90.9% via iterative auditing.

#factuality#benchmarks#llm-agents
📄
ArXiv AI173d ago

CliqueFlowmer Boosts Offline Materials Optimization

CliqueFlowmer introduces an offline model-based optimization approach for computational materials discovery, fusing direct target property optimization into generation. It leverages clique-based MBO within transformer and flow architectures, outperforming generative baselines. The code is open-sourced on GitHub for broader use.

#materials-discovery#open-source
📄
ArXiv AI173d ago

AI Agents Evaluate Product Concepts

This arXiv paper proposes an LLM-based multi-agent system to automate product concept evaluation, addressing biases in traditional methods. Eight specialized agents assess technical and market feasibility using RAG and real-time search. A case study on display monitors shows results matching senior experts.

#multi-agent#product-evaluation#llm-tools
📄
ArXiv AI173d ago

Agentic LLM Planning via PDDL Simulation

PyPDDLEngine is an open-source PDDL simulation engine enabling LLMs to plan interactively via tool calls. It allows step-wise action selection, state observation, and resets. Agentic LLM planning achieves 66.7% success on IPC Blocksworld, edging direct prompting at 63.7% but trailing classical planners at 85.3%.

#agentic-planning#pddl-simulation#llm-benchmarks
📄
ArXiv AI173d ago

Agentic AI Framework for Real-Time Services

This arXiv paper introduces a framework for real-time AI services across device-edge-cloud continuum using DAG-modeled service dependencies for decentralized resource allocation. Hierarchical DAGs enable stable pricing and optimal allocations, while complex ones need hybrid architectures to reduce volatility. An ablation study confirms topology's role in stability, efficiency, and governance trade-offs.

#agentic-computing#resource-allocation#dag-topology
📄
ArXiv AI173d ago

Agentic AI Powers Conversational Demand Response

Introduces CDR, enabling bidirectional natural language coordination between energy aggregators and prosumers via agentic AI. Features a two-tier multi-agent architecture with aggregator dispatching and prosumer HEMS optimization. Open-source release includes prompts, logic, and simulations, with interactions under 12 seconds.

#agentic-ai#demand-response#multi-agent
🐯
虎嗅173d ago

Uni Cuts 16 Majors for AI Education Shift

China Media University axes 16 majors like translation and photography, citing AI's takeover of routine skills in the human-machine era. It urges refocusing on irreplaceable human strengths like creativity and ethics. Teachers evolve into AI coaches and learning designers.

#education-reform#ai-impact#higher-education
🐯
虎嗅173d ago

Qwen Lead Lin Junyang Quits Alibaba

Alibaba's youngest P10, Lin Junyang, resigns after Qwen3.5 success and Musk praise, amid company pushback on personal influence. Highlights 'big carrot small carrot' dynamic where firms control talent placement. Stresses need for obedient high-performers over independent leaders.

#leadership-change#open-source#talent-dynamics
🐯
虎嗅173d ago

Benchmarking AI Toward Digital Scientists

Science explores AI scientific smarts beyond memorization via GPQA, where o1 scores 80%+ vs. experts' 65-70%. Shifts to process audits and closed-loop lab tests for true reasoning. Humans lead in paradigm shifts despite AI optimization prowess.

#benchmarks#scientific-ai#reasoning-eval
🗾
ITmedia AI+ (日本)173d ago

X Adds Grok Image Edit Block Option

X now lets users select to block Grok AI edits when posting images. The feature is rolling out to select users as of March 9. It specifically blocks edits triggered by mentioning Grok's official account.

#user-control#ai-optout#social-ai
🗾
ITmedia AI+ (日本)173d ago

HubSpot Survey: AI Transforms Sales Roles

Buyers' AI adoption is reducing demand for traditional proposal-focused salespeople. HubSpot's latest survey examines optimal 'sales x AI' strategies. It highlights evolving human roles in AI-driven sales processes.

#sales-ai#crm-strategy#buyer-trends
📊
Bloomberg Technology173d ago

OpenClaw Stocks Surge on Shenzhen Policy Boost

Chinese companies linked to open-source AI agent software OpenClaw saw stock advances. This follows Shenzhen authorities' measures to support development of OpenClaw-based tools. The move signals growing policy-backed adoption in China.

#china-policy#ai-agents#stock-impact
🇦🇺
iTNews Australia173d ago

Saviynt roundtable on AI agents security

iTnews hosted a roundtable lunch in Sydney on securing AI agents and Non-Human Identities (NHIs) with Saviynt. Photos from the event at Establishment are featured. Discussion focused on security challenges for AI deployments.

#ai-security#nhi#roundtable
Page 1711 of 1963