All Updates
Page 1711 of 1963
March 9, 2026
RoboLayout: Agent-Friendly 3D Scenes
RoboLayout extends LayoutVLM with agent-aware reasoning and reachability constraints for generating navigable 3D indoor scenes from language. It supports diverse agents like robots, humans, and animals via abstract representations. Local refinement optimizes problematic placements efficiently while preserving semantic coherence.
RL for Climate-Resilient Transport
Proposes a reinforcement learning (RL) framework for long-term flood adaptation in urban transport systems. Integrates rainfall projections, flood modeling, transport simulation, and impact quantification as an integrated assessment model (IAM). Outperforms traditional optimization in Copenhagen case study by discovering coordinated strategies under climate uncertainty.
Reasoning Models Fail CoT Control
New CoT-Control suite reveals reasoning models have low controllability over chain-of-thought vs. final outputs, e.g., Claude Sonnet 4.5 at 2.7% CoT control but 61.9% output. Controllability drops with larger models, RL training, compute, and difficulty. Results suggest CoT monitoring remains viable despite evasion attempts.
ProEvolve: Programmable Agent Benchmark Evolution
ProEvolve is a graph-based framework for scalable, controllable evolution of agent environments to test LLM agents' adaptability to real-world changes. It uses a typed relational graph to unify data, tools, and schemas, enabling updates via graph transformations that propagate coherently. The system generates 200 environments and 3,000 task sandboxes for benchmarking.
New Aggregative Semantics for QBAF
Researchers propose aggregative semantics for Quantitative Bipolar Argumentation Frameworks (QBAF), aggregating attackers and supporters separately unlike modular approaches. This involves a three-stage process: computing global attacker and supporter weights before combining with intrinsic argument weights. The work explores properties, principles, and compares 500 variants for diverse behaviors.
Mistral Launches Local Japanese Speech AI Voxtral 2
Mistral AI announced Voxtral Transcribe 2, a speech recognition model that runs locally and supports Japanese. It includes two variants: one for high-accuracy, low-cost batch processing and another for ultra-low latency real-time transcription.
MACRO: Self-Evolving Medical Imaging Agent
MACRO is a self-evolving agent that autonomously discovers reusable composite tools from verified execution trajectories in medical imaging tasks. It grounds tool selection in visual-clinical context via image-feature memory and reinforces reliable use through GRPO-like training. Experiments show superior multi-step accuracy and cross-domain generalization over baselines.
DeepFact: Co-Evolving Benchmarks for LLM Factuality
DeepFact proposes Audit-then-Score (AtS) to create evolving benchmarks for verifying factuality in deep research reports from search-augmented LLM agents. It introduces DeepFact-Bench, a versioned benchmark with auditable rationales, and DeepFact-Eval, a verification agent that outperforms existing tools. Expert accuracy on claims rises from 60.8% to 90.9% via iterative auditing.
CliqueFlowmer Boosts Offline Materials Optimization
CliqueFlowmer introduces an offline model-based optimization approach for computational materials discovery, fusing direct target property optimization into generation. It leverages clique-based MBO within transformer and flow architectures, outperforming generative baselines. The code is open-sourced on GitHub for broader use.
AI Agents Evaluate Product Concepts
This arXiv paper proposes an LLM-based multi-agent system to automate product concept evaluation, addressing biases in traditional methods. Eight specialized agents assess technical and market feasibility using RAG and real-time search. A case study on display monitors shows results matching senior experts.
Agentic LLM Planning via PDDL Simulation
PyPDDLEngine is an open-source PDDL simulation engine enabling LLMs to plan interactively via tool calls. It allows step-wise action selection, state observation, and resets. Agentic LLM planning achieves 66.7% success on IPC Blocksworld, edging direct prompting at 63.7% but trailing classical planners at 85.3%.
Agentic AI Framework for Real-Time Services
This arXiv paper introduces a framework for real-time AI services across device-edge-cloud continuum using DAG-modeled service dependencies for decentralized resource allocation. Hierarchical DAGs enable stable pricing and optimal allocations, while complex ones need hybrid architectures to reduce volatility. An ablation study confirms topology's role in stability, efficiency, and governance trade-offs.
Agentic AI Powers Conversational Demand Response
Introduces CDR, enabling bidirectional natural language coordination between energy aggregators and prosumers via agentic AI. Features a two-tier multi-agent architecture with aggregator dispatching and prosumer HEMS optimization. Open-source release includes prompts, logic, and simulations, with interactions under 12 seconds.
Uni Cuts 16 Majors for AI Education Shift
China Media University axes 16 majors like translation and photography, citing AI's takeover of routine skills in the human-machine era. It urges refocusing on irreplaceable human strengths like creativity and ethics. Teachers evolve into AI coaches and learning designers.
Qwen Lead Lin Junyang Quits Alibaba
Alibaba's youngest P10, Lin Junyang, resigns after Qwen3.5 success and Musk praise, amid company pushback on personal influence. Highlights 'big carrot small carrot' dynamic where firms control talent placement. Stresses need for obedient high-performers over independent leaders.
Benchmarking AI Toward Digital Scientists
Science explores AI scientific smarts beyond memorization via GPQA, where o1 scores 80%+ vs. experts' 65-70%. Shifts to process audits and closed-loop lab tests for true reasoning. Humans lead in paradigm shifts despite AI optimization prowess.
X Adds Grok Image Edit Block Option
X now lets users select to block Grok AI edits when posting images. The feature is rolling out to select users as of March 9. It specifically blocks edits triggered by mentioning Grok's official account.
HubSpot Survey: AI Transforms Sales Roles
Buyers' AI adoption is reducing demand for traditional proposal-focused salespeople. HubSpot's latest survey examines optimal 'sales x AI' strategies. It highlights evolving human roles in AI-driven sales processes.
OpenClaw Stocks Surge on Shenzhen Policy Boost
Chinese companies linked to open-source AI agent software OpenClaw saw stock advances. This follows Shenzhen authorities' measures to support development of OpenClaw-based tools. The move signals growing policy-backed adoption in China.
Saviynt roundtable on AI agents security
iTnews hosted a roundtable lunch in Sydney on securing AI agents and Non-Human Identities (NHIs) with Saviynt. Photos from the event at Establishment are featured. Discussion focused on security challenges for AI deployments.