All Updates
Page 802 of 1665
April 29, 2026
Baidu GenFlow 4.0 Handles Office, AI Agents
Baidu has launched GenFlow 4.0, fully covering the Office suite (Word, Excel, PowerPoint). It also supports managing 'cow horse shrimp' AI agents, claiming to finish a week's work in 10 minutes.
AI Short Dramas Conquer Hengdian by 2026
Hengdian studios fade as AI short dramas explode on Douyin, hitting 1-1.2B daily revenue vs human dramas' halving. Platforms use dual-track: human for premium, AI for volume with superior ROI. Long-video sites like iQiyi now test AI mid-length series.
Cloud Leader Launches Longxia, Altman Joins
China's top cloud provider released the 'Longxia' (Lobster) edition. Sam Altman attended the cloud event despite his lawsuit, with GPT-5.5 support arriving soon.
Knowledge Tools Vital in Agent Era?
This article questions if products centered on knowledge accumulation hold irreplaceable value in the AI Agent era. It reviews the latest evaluation of Tencent's IMA-Copilot mode, evolving from 'lobster' to 'knowledge brain.' It explores what remains essential once Agents perform tasks independently.
StoryTR: ToM for Narrative Video Retrieval
StoryTR introduces the first video moment retrieval benchmark requiring Theory of Mind (ToM) reasoning for narrative understanding, with 8.1k samples from short-form videos. Current models like Gemini-3.0-Pro struggle at 0.53 Avg IoU due to lacking intent and causality inference. A 7B Shorts-Moment model, trained via Agentic Data Pipeline on ToM chains, achieves +15.1% relative IoU improvement.
SoccerRef-Agents: AI Soccer Refereeing Agents
SoccerRef-Agents is a multi-agent framework for automated soccer refereeing, integrating MLLMs with referee expertise via cross-modal RAG. It introduces SoccerRefBench benchmark with 1,200 theory questions and 600 foul videos, plus RefKnowledgeDB from Laws of the Game. The system outperforms general MLLMs in accuracy and explanations, with code and data to be released.
OpenAI's Five-Part AI Cyber Defense Plan
OpenAI has unveiled a five-part action plan to enhance cybersecurity in the Intelligence Age. It emphasizes democratizing AI-powered cyber defense tools for broader access. The focus is on protecting critical systems from advanced threats.
LLMs Learn Safety from 1-Bit Signals
EPO-Safe enables LLMs to discover hidden safety objectives using only binary danger warnings per timestep, without observing the true reward function. Through iterative action planning and reflection, agents evolve human-readable safety specifications in 1-2 rounds on gridworlds and text scenarios. It outperforms reward-driven reflection, which promotes reward hacking, and remains robust to noisy signals.
LEGO: LLM Skill-Based EDA Platform
LEGO is a unified platform decomposing digital front-end design into six steps with 42 composable circuit skills from open-source projects. It automates skill extraction and enables fast retrieval via Agent Skill RAG. Evaluations on hard VerilogEval v2 problems show 80.5% Pass@1, outperforming baselines like GPT-5.2-Codex.
GSAR: Typed Grounding for LLM Hallucinations
GSAR is a new framework for hallucination detection and recovery in multi-agent LLMs, partitioning claims into grounded, ungrounded, contradicted, and complementary types. It uses evidence-type-specific weights and an asymmetric contradiction-penalized score, coupled with a three-tier decision function (proceed, regenerate, replan) under a compute budget. Evaluations on FEVER dataset with multiple LLM judges show strong performance gains.
CAP-CoT Boosts LLM CoT Stability
CAP-CoT introduces cycle adversarial prompting to enhance Chain-of-Thought (CoT) reasoning accuracy and stability in LLMs. A forward solver generates chains, an adversarial challenger creates flawed ones with targeted errors, and a feedback agent provides structured corrections to iteratively update prompts. Experiments on six benchmarks and four LLMs show reduced variability and improved robustness after 2-3 cycles.
Bias Mitigation Evaluated in LLM Judges
A study evaluates nine debiasing strategies across five LLM judges from Google, Anthropic, OpenAI, and Meta on three benchmarks. Style bias dominates (0.76-0.92), while models correctly distinguish quality from length. Combined budget debiasing improves Claude Sonnet 4 by +11.2 pp, with code and dataset released on GitHub.
Analyzing Reasoning Shortcuts in Neurosymbolic AI
Neurosymbolic systems often learn reasoning shortcuts that satisfy constraints but miss intended concepts. Researchers formalize this as a constraint satisfaction problem, proving conditions for unique mappings and developing an ASP-based algorithm to verify and repair shortcuts. Experiments across benchmarks validate the approach with complexity and sample bounds.
AI Identity Gaps for Agents
This arXiv paper defines AI Identity as the alignment between an agent's declared purpose and observed actions, bounded by confidence. It compares human and AI identity across substrate, persistence, verifiability, and legal standing, revealing fundamental asymmetries. The analysis identifies five critical gaps in standards for autonomous agents, urging foundational research.
AdaPlan-H: Adaptive Hierarchical LLM Planning
AdaPlan-H proposes self-adaptive hierarchical planning for LLM agents, starting with coarse macro plans and refining based on task complexity. Inspired by cognitive progressive refinement, it balances detail levels for optimal performance. Experiments show higher success rates, reduced overplanning, with code releasing on GitHub.
AdaMamba Revolutionizes Long-Term Time Series Forecasting
AdaMamba is a novel framework that embeds adaptive, context-aware frequency analysis into Mamba's state-space updates for superior long-term time series forecasting (LTSF). It features an interactive patch encoding module for inter-variable dynamics and an adaptive frequency-gated module with unified time-frequency forgetting gates. The model outperforms state-of-the-art methods on seven benchmarks while maintaining efficiency, with code available on GitHub.
Active Inference Phenotypes AI Agency
Researchers propose a minimal definition of AI agency based on intentionality, rationality, and explainability, modeled as a variational POMDP minimizing expected free energy. Using a T-maze paradigm, they demonstrate how empowerment (channel capacity between actions and observations) distinguishes zero-, intermediate-, and high-agency phenotypes. The work suggests shifting AI governance from external constraints to internal modulation of prior preferences.
StepStars Releases Step Image Edit 2
Tierstep Stars launched the new generation image generation and editing model Step Image Edit 2. It's now fully available on the Tierstep Stars open platform and Step Plan.
YoooClaw CΒ·ONE: Portable AI Hardware Gateway
ε°ζ°ζ΄Ύ invites crowdtesting of YoooClaw CΒ·ONE, a portable AI hardware entry point. AI is evolving from mere question-answering to proactive task completion with context awareness. Users can explore this device to experience hands-on AI integration.
Palmread Accelerates AI Short Drama by 2026
Palmread Tech plans full AI short drama industrialization in 2026 for scaled content ecosystem. Will boost overseas push with digital reading, premium shorts, AI comics. Domestic costs to shift significantly.