All Updates

Page 594 of 1669

May 16, 2026

๐Ÿ“„
ArXiv AIโ€ข77d ago

SkillFlow: Flow-Driven Recursive Skill Evolution for Agentic Orchestration

SkillFlow is a new framework that uses flow-matching and Tempered Trajectory Balance to automate task orchestration in agentic systems. It enables autonomous skill evolution by identifying decision gaps and creating or pruning skills based on principled training signals.

#agentic-workflows#flow-matching#autonomous-agents
๐Ÿ“„
ArXiv AIโ€ข77d ago

MSIFR: Cutting LLM Synthetic Data Costs by Up to 78%

MSIFR is a new training-free framework that terminates low-quality LLM generation trajectories early. By using rule-based validators at intermediate stages, it significantly reduces token waste without requiring architectural changes or additional training.

#synthetic-data#cost-optimization#llm-efficiency
๐Ÿ“„
ArXiv AIโ€ข77d ago

Modeling Bounded Rationality in Pharmacist Decision-Making

This research introduces an attention-guided framework that mimics how pharmacists prioritize drug shortages under high-stakes conditions. By dynamically decomposing tasks into high-cost reasoning and low-cost monitoring, the model achieves stable performance without requiring complete state reasoning.

#agentic-ai#cognitive-modeling#decision-making
๐Ÿ“„
ArXiv AIโ€ข77d ago

MathAtlas: New Benchmark for Graduate-Level Autoformalization

MathAtlas is a large-scale benchmark featuring 52k graduate-level mathematics entries, including theorems and proofs. It introduces a dependency graph to test the reasoning capabilities of models on complex, multi-step mathematical structures.

#autoformalization#mathematics#benchmarking
๐Ÿ“„
ArXiv AIโ€ข77d ago

LLM Agents Outperform Heuristics in Combinatorial Optimization

This research introduces a framework for using LLM agents to synthesize specialized solver code based on task distribution samples. The synthesized solvers significantly outperform standard heuristics and industry-standard tools like Gurobi in both execution speed and solution quality.

#llm-agents#algorithm-design
๐Ÿ“„
ArXiv AIโ€ข77d ago

ClawForge: Benchmarking Command-Line Agents Under State Conflict

ClawForge is a new benchmark framework designed to evaluate command-line agents by simulating realistic, persistent state conflicts. It tests how agents handle partial or stale artifacts, revealing that current frontier models struggle with state inspection before taking action.

#agentic-workflow#benchmarking
๐Ÿ“„
ArXiv AIโ€ข77d ago

ChromaFlow: Negative Ablation Study on Agent Orchestration Overhead

This research introduces ChromaFlow, a framework for evaluating autonomous agents, and reveals that increased orchestration complexity often leads to higher operational noise without improving performance. The study highlights that simpler, more deterministic approaches may be more effective for reliable agent evaluation.

#autonomous-agents#orchestration#evaluation-framework
๐Ÿ“„
ArXiv AIโ€ข77d ago

Bridging Legal Interpretation and Formal Logic for AI

This research proposes a neuro-symbolic approach to address the reliability issues of LLMs in legal practice. By combining large language models with formal verification, the framework aims to ensure that legal reasoning remains grounded in source texts.

#neuro-symbolic#legal-tech#formal-verification
๐Ÿ“„
ArXiv AIโ€ข77d ago

Boosting Weak Reasoning Models with Agentic Committees

This research demonstrates that orchestrating a committee of weak reasoning models can match the performance of top-tier models like Gemini 3 Pro. By using verifier-backed selection, the system effectively identifies correct solutions from a pool of weak model outputs.

#agentic-workflows#reasoning-models#model-orchestration
๐Ÿ“„
ArXiv AIโ€ข77d ago

AI Benchmarking is Fragmented and Narrative-Driven

Researchers analyzed 231 benchmarks from 11 major AI builders in 2025, finding that most benchmarks are used only once and serve as marketing tools rather than standardized scientific metrics. The study highlights a lack of cross-model comparability and a tendency to prioritize market positioning over rigorous evaluation.

#benchmarking#model-evaluation#agi-claims
๐Ÿฏ
่™Žๅ—…โ€ข77d ago

The Data Bottleneck in Embodied AI and Robotics

The article explores the 'data desert' in robotics, highlighting the difficulty of acquiring multi-modal sensor data compared to LLMs. It details the four-layer data pyramid, emphasizing the high cost and labor-intensive nature of high-quality teleoperation data.

#embodied-ai#robotics#data-collection
๐Ÿฏ
่™Žๅ—…โ€ข77d ago

AI Agent Successfully Earns Money via Bug Bounties

An AI agent (Codex) successfully completed a software security audit task, earning $16.88 in bounty. This experiment demonstrates the potential for AI agents to participate in real-world economic systems through defined software tasks.

#ai-agent#bounty-program#automation
๐Ÿ”ฅ
36ๆฐชโ€ข77d ago

Starbucks Opens First Tech-Focused Office in India

Starbucks is establishing its first dedicated corporate office in India specifically for recruiting technical talent. This move signals a strategic shift toward strengthening the company's internal digital and technical infrastructure.

#tech-hiring#retail-tech
๐Ÿฏ
่™Žๅ—…โ€ข77d ago

Humanoid Robots: From Performance Art to Industrial Reality

The article analyzes the current state of humanoid robotics, noting that while entertainment and demos are common, real-world industrial application remains the primary hurdle. Logistics and tourism have emerged as the first viable sectors for deployment, while complex manufacturing tasks still face significant technical barriers.

#embodied-ai#humanoid-robotics#supply-chain
๐Ÿ”ฅ
36ๆฐชโ€ข77d ago

China Regulators Probe Pharmacies for Medical Insurance Fraud

The National Healthcare Security Administration (NHSA) has summoned four major pharmacy chains for investigation following reports of illegal medical insurance fund usage. The pharmacies were found swapping medical insurance funds for non-medical items like cosmetics.

#fraud-detection#regulatory-tech
๐Ÿ 
ITไน‹ๅฎถโ€ข77d ago

Microsoft restricts Claude Code usage for internal teams

Microsoft is phasing out Claude Code licenses for internal engineering teams by the end of June, mandating a transition to GitHub Copilot CLI. The move aims to better integrate development workflows with Microsoft's internal security and infrastructure standards.

#enterprise-ai#dev-tools#microsoft-strategy
๐Ÿ”ฅ
36ๆฐชโ€ข77d ago

OpenAI acquires voice cloning startup Weights.gg

OpenAI has quietly acquired Weights.gg, a startup specializing in AI-powered voice cloning tools. The deal includes the acquisition of the team and all intellectual property, following the startup's service shutdown in March.

#voice-cloning#m-and-a#audio-synthesis
๐Ÿฏ
่™Žๅ—…โ€ข77d ago

The trend of 'Cyber-Diagnosis' and mental health labeling

Social media users are increasingly using clinical labels like ADHD and ASD as social currency or self-explanation tools. Experts warn that this 'cyber-diagnosis' trend risks trivializing genuine conditions and creating self-limiting narratives for young people.

#mental-health#social-media#psychology
๐Ÿฏ
่™Žๅ—…โ€ข77d ago

Airlines vs OTA: The Value of Information Aggregation

The article analyzes the intensifying conflict between airlines and OTAs, arguing that OTAs provide essential value in information aggregation and service experience that airlines struggle to replicate due to organizational and cost constraints.

#platform-economy#business-strategy
๐Ÿฏ
่™Žๅ—…โ€ข77d ago

Heartbeat Interactive's New Game Launch After Acquisition

Heartbeat Interactive (ๅฟƒๅ…‰ๆต็พŽ) has launched its new strategy RPG 'Radium Flash' following its acquisition by ByteDance's Nuverse. The studio aims to leverage its 3D art expertise and unique 'tower-attack' gameplay to overcome past financial and organizational challenges.

#gaming-industry#m-and-a#game-design
Page 594 of 1669