All Updates
Page 647 of 1678
May 13, 2026
Tencent Seeks 2026 Win After Setbacks
Tencent has faced two consecutive stumbles in performance. Executive Yao Shunyu is now steadying CEO Ma Huateng amid challenges. The company urgently needs a major victory in 2026 to regain momentum.
47 Open-Ended Tasks Redefine Agent Benchmarks
In the Auto Research era, 47 tasks without standard answers have become the essential benchmark for evaluating AI Agent capabilities. This list marks the official entry into the 'iteration optimization' era, focusing on Agents' ability to handle unstructured problems.
Funds Add to Electronics Stocks, Sell Finance Midday
Morning main funds net inflow to electronics, power equipment, machinery, utilities; outflow from non-bank finance, food & beverage, biopharma. Top inflows: Industrial Fulian RMB5.192B, Tianfu Comm RMB2.147B, Shengyi Tech RMB1.935B. Top outflows: Guangxun Tech RMB1.273B, Zhongji Innolight RMB1.238B, Huagong Tech RMB1.102B.
Family Sues OpenAI Over ChatGPT Drug Overdose
A family is suing OpenAI, claiming that ChatGPT provided drug use advice to Sam Nelson, leading to his accidental overdose. The complaint alleges this began with the launch of GPT-4o. This marks a significant liability case against AI chatbots.
SK Hynix Profit Share Fuels Samsung Strike
SK Hynix locked in 10% annual profit sharing with unions for 10 years, yielding massive AI-driven HBM bonuses. Samsung rejects similar formula, faces historic strike from May 21 impacting 3-4% global DRAM. Signals profit redistribution push in AI supply chain.
Google Releases Android AI System, Realizing Apple's Vision
Google has unveiled its new Android AI system, integrating Gemini models more deeply into the mobile OS. The update aims to provide a unified, multi-modal AI experience across the Android ecosystem.
Self-Generated Data Boosts LLM Reinforcement Learning Performance
This research introduces a method using diverse, self-generated data based on George Polya's problem-solving strategies to improve LLM performance during mid-training. The study demonstrates that models trained with these multiple reasoning approaches achieve superior results in mathematical, coding, and narrative tasks after RL.
Personality Drives AI Agents' Social Behavior
Researchers deployed 13 OpenClaw AI agents on Moltbook, a Reddit-like social network, varying personality specs, LLM models, and rules. Personality was the strongest factor influencing response length, while models and rules moderately affected rhetorical style and topic engagement. The study provides empirical guidance for designing agents in social environments.
OracleTSC Stabilizes LLM-Based Traffic Signal Control Systems
OracleTSC introduces a reward hurdle and uncertainty regularization to stabilize reinforcement learning for traffic signal control. This approach allows compact models like LLaMA3-8B to significantly improve traffic efficiency and cross-intersection generalization.
LLM-Guided Co-Training Excels in Low-Label Crisis Classification
First empirical evaluation of LLM-guided semi-supervised methods for crisis tweet classification. LG-CoTrain outperforms baselines with 5-25 labels per class, achieving top Macro F1. Compact models beat zero-shot LLMs for practical disaster response.
Latent Personality Alignment: Improving Harmlessness Without Harmful Data
Latent Personality Alignment (LPA) is a new, sample-efficient defense method that aligns LLMs using abstract personality traits instead of massive datasets of harmful examples. It achieves superior robustness and generalization to unseen attacks while maintaining high model utility.
Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care
A study on MedSyn shows that interactive LLM support significantly improves diagnostic accuracy for emergency physicians. Residents experienced the most notable gains in diagnostic correctness when using the system as an iterative query tool.
CODS 2025 AssetOpsBench Challenge Retrospective Analysis
This retrospective analyzes the CODS 2025 AssetOpsBench challenge, revealing that public leaderboards often fail to predict real-world execution performance. The study highlights that successful systems prioritize robust guardrails over complex agent architectures.
Biologically-Inspired Memory Architecture for LLM Agents
This research introduces a memory architecture for LLM agents that mimics human cognitive processes, including sleep-phase consolidation and interference-based forgetting. It effectively manages long-term interaction data while maintaining high retrieval accuracy and reducing storage requirements.
Benchmarking AI Reliability in Healthcare
New arXiv paper argues current AI benchmarks fail to measure reliability, safety, and clinical relevance in real-world healthcare workflows. Frontier models ace exams but score low (0.53-0.85) on documentation, decision support, and admin tasks. Calls for principled frameworks to bridge performance-utility gap.
Anchored Bipolicy Self-Play Boosts AI Safety
Researchers propose Anchored Bipolicy Self-Play to address limitations in standard self-play red teaming for AI safety. It uses distinct role-specific LoRA adapters on a frozen base model, avoiding self-consistency collapse and maintaining adversarial pressure. Evaluations on Qwen2.5 models show superior safety and efficiency over baselines.
Alignment as Jurisprudence
This essay compares AI alignment and jurisprudence, both using language to guide powerful decision-makers like judges or AIs. It draws parallels with Dworkin's interpretivism, Sunstein's analogical reasoning, Constitutional AI, and case-based reasoning. It advocates legally-inspired finetuning for alignment and AI's role in improving law.
AI-Care: Agentic Task Coordination for Alzheimer's Care
AI-Care is a conversational agentic system designed to assist individuals with Alzheimer's disease in managing daily tasks like calendar scheduling. It utilizes a stateful orchestration approach to ensure safety and reliability in caregiving environments.
AI Agent Decisions Lack Trust, Need Human Checks
Dynatrace's survey shows AI agent decision-making is not yet trusted and requires human verification as a prerequisite. Enterprises mainly use human oversight methods for validation. The article details common human-led verification approaches.
2000+ Scientists Warn NSF Cuts Cede AI Lead to China
Over 2000 scientists protest Trump admin dissolving NSF's NSB, fearing weakened US science edge amid China rivalry. NSF faces 57% budget slash; China matches/surpasses US R&D, leads Nature index in AI/semiconductors. Policy shifts prioritize national security over broad research.