All Updates
Page 1699 of 1965
March 10, 2026
OpenClaw Craze: FOMO and Risks
OpenClaw sparks absurd install frenzy in China driven by FOMO, not utility; security nightmares per Cisco Talos with plaintext tokens, malicious plugins. Users burn tokens fast without real gains; vendors profit via services/courses. Critiques it as social signaling over productivity.
Scam-Label Firm Rises as India's DeepSeek
A company derided as a scam has emerged as India's DeepSeek equivalent in 27 months. The turnaround story originates from December 2023.
OpenClaw Boom Fuels Risky Install Services
OpenClaw, dubbed 'lobster' in China, sparks paid installation services on platforms like Taobao, generating 30-45万 RMB monthly. Despite hype, it faces high setup barriers, security flaws in 42k+ instances, and API burn risks. Alibaba execs warn of AI over-reliance.
OpenClaw Virals, Saves Chinese LLMs
OpenClaw has unexpectedly gone viral in popularity. This provides relief to Chinese AI firms Zhipu, MiniMax, and Kimi. The piece suggests AI's killer app may be omnipresent APIs, not traditional apps.
AI Stocks Mint Employee Millionaires
Zhipu AI and MiniMax stocks surged 500%+, making half of Zhipu employees shareholders with avg 36M RMB paper wealth. Short cycles and hot capital enable broad equity without dilution. But high burn and volatility question sustainability.
China Gen Z Trusts Chatbots to Move Markets
China's Gen Z day traders increasingly rely on chatbots for trading decisions, enabling them to influence financial markets. This young cohort is emerging as a key driver of investment growth in the world's second-largest economy.
SymLang: AI Equation Discovery from Noisy Data
SymLang is a new framework integrating symmetry-constrained grammars, language-model-guided program synthesis, and Bayesian model selection to discover governing equations from noisy, partial observations. It achieves 83.7% exact structural recovery across 133 dynamical systems under 10% noise, outperforming baselines by 22.4 points. The open-source tool reduces extrapolation errors by 61% and minimizes conservation violations.
RL Agents Enhance Option Hedging
Introduces RLOP and QLBS, two RL frameworks prioritizing shortfall probability for better option hedging. Evaluated on SPY and XOP options using realized hedging outcomes and tail risks. RLOP reduces shortfalls and excels in stress scenarios for autonomous derivatives management.
Open-Source StarCraft II RL Benchmark Launch
研究團隊推出 Two-Bridge Map Suite,這是 StarCraft II 的獨立開源基準測試,填補完整遊戲與小遊戲之間的複雜度空白。它透過停用資源收集、基地建造與戰爭迷霧,聚焦長距離導航與微操戰鬥。作為輕量 Gym 相容 PySC2 包裝器發布,鼓勵廣泛採用為標準基準。
MultiGen: Editable Multiplayer Diffusion Worlds
MultiGen introduces explicit external memory for user-controlled, editable environments in diffusion game engines. It decomposes generation into Memory, Observation, and Dynamics modules for reproducible experiences. This enables real-time multiplayer rollouts with coherent viewpoints and cross-player interactions.
MARL Boundary Drift Causes Continual Learning Woes
Reusable RL decision structures depend on agent-world boundary. In multi-agent settings, peer policy updates create non-stationary MDPs, shrinking invariant state-action cores. Frames this as endogenous continual RL challenge from boundary instability.
LLM Response Length Shapes Critical Thinking
A study examines how LLM response length affects users' accuracy in evaluating LLM-generated reasoning on critical thinking tasks. LLM correctness strongly boosts user accuracy, moderated by length: medium-length explanations optimize performance when LLM is incorrect. Findings suggest design opportunities for LLM systems to enhance transparent reasoning.
LieCraft Tests LLM Deception in Multi-Agent Games
LieCraft is a new multi-agent framework and sandbox for evaluating deception in LLMs via a multiplayer hidden-role game. Players adopt cooperator or defector roles in 10 real-world scenarios like childcare and hospital allocation. Tests on 12 top LLMs show all models defect, deceive, and lie to achieve goals.
LEAD Breaks LLM No-Recovery Bottleneck
Long-horizon LLM execution is unstable due to a 'no-recovery bottleneck' from extreme decomposition and non-uniform errors. LEAD uses lookahead-enhanced atomic decomposition with short-horizon validation and overlapping rollouts to enable error correction. This allows o4-mini to solve Checkers Jumping puzzles up to n=13, surpassing n=11 limits.
DLI for Trustworthy Agentic Web
The paper proposes a Distributed Legal Infrastructure (DLI) to govern the agentic web, where AI agents act autonomously. It outlines five layers: self-sovereign agent identities, cognitive constraints, decentralized adjudication, market regulation, and portable frameworks. This enables legality, accountability, and interoperability across AI systems.
Context Specs for Relevant AI Evaluations
Organizations struggle to gain value from AI deployments due to inadequate evaluation methods that ignore operational realities. The paper introduces 'context specification,' a process to translate stakeholder perspectives into explicit, measurable constructs of properties, behaviors, and outcomes. This roadmap helps evaluate AI systems in real deployment contexts for better decision-making.
Best-of-Tails: Adaptive LLM Alignment
Best-of-Tails (BoT) addresses the optimism-pessimism dilemma in inference-time LLM alignment by adaptively selecting responses based on reward distribution tails. It uses the Hill estimator for per-prompt tail heaviness and Tsallis divergence as a tunable regularizer. BoT outperforms baselines in math, reasoning, and human-preference tasks.
Anthropic Adds Plugins to Cowork for Custom Claude
Anthropic added a plugin feature to its AI agent Cowork, allowing customization of business-specific Claude agents. Users can combine skills, connectors, slash commands, and sub-agents. Enhances task-oriented AI workflows.
AceMAD Breaks Martingale Curse in MAD
Multi-Agent Debate (MAD) suffers from the Martingale Curse, where agents converge to erroneous consensus due to correlated errors. AceMAD breaks this by using asymmetric cognitive potential energy via peer-prediction, directing convergence toward truth. It outperforms baselines on challenging subsets across six benchmarks.
Baidu Resumes Driverless Ops in UAE Cities
Baidu's Luobo Kuai Run has resumed fully unmanned testing and operations in Dubai and Abu Dhabi. The simultaneous rollout in both cities speeds up overseas driverless deployment. This sets the stage for scaled commercial operations in the UAE.