All Updates

Page 832 of 1664

April 27, 2026

๐Ÿ‡ญ๐Ÿ‡ฐ
SCMP Technologyโ€ข89d ago

China's Physical AI Hits Roads, Skies, Factories

Physical AI, merging machines with AI for environmental interaction, is surging in China. Smarter robots, drones, and driverless cars appear on roads, factory floors, skies, and stages. Delivery drones fly over Shenzhen, bots ride subways, and autonomous vehicles operate on public roads.

#embodied-ai#robotics#drones
๐Ÿ’ฐ
้’›ๅช’ไฝ“โ€ข89d ago

Entrust Enterprise Security to Agents?

Questions if enterprises can delegate security to AI agents. Examines human-machine collaboration in security industry. Provides ToB industry observations.

#cybersecurity#human-ai-collab#tob-security
๐Ÿ”ฅ
36ๆฐชโ€ข89d ago

QClaw v0.2.14 Adds Hermes & DeepSeek V4

Tencent's QClaw released v0.2.14, adding support for Hermes framework to create and run Agents. Features model switching to Hy3 preview and DeepSeek-V4Pro, new connectors for Baidu Netdisk, Ctrip, Fliggy, Tencent News, plus WeChat voice/file sharing and Tencent Docs collaboration.

#agent-framework#multi-llm#connectors
๐Ÿ’ฐ
้’›ๅช’ไฝ“โ€ข89d ago

Geely-Backed Qianli's AI Valuation Questioned

Qianli Technology, supported by Geely, has seen its valuation surge due to AI hype. The article scrutinizes the actual substance behind this valuation boost. Meanwhile, Huaqin Technology and Shenghong Technology complete A+H stock listings on Hong Kong exchange.

#valuation-hype#auto-ai#stock-listing
๐Ÿ 
ITไน‹ๅฎถโ€ข89d ago

Xiaomi Chip-AI-OS Convergence Schedule Confirmed

Xiaomi confirms schedule for next-gen self-developed chip, AI large model, and OS converging in a terminal product, slightly later than rumors. Xuanjie O1 chip has shipped over 1 million units and will cover phone, tablet, car, wearables ecosystems. Lei Jun eyes 2026 for full self-chip/OS/AI integration in one device.

#self-developed-chip#ai-model#os-integration
๐Ÿ“Š
Bloomberg Technologyโ€ข89d ago

DeepSeek Slashes Fees on New Flagship Model

DeepSeek has launched its new flagship AI model with aggressively low-priced plans. This pricing strategy intensifies competition in China's AI sector. The move aims to challenge Silicon Valley's leading models.

#china-ai#price-war#competition
๐Ÿ”ฅ
36ๆฐชโ€ข89d ago

Hang Seng Tech Up 1.31% Midday Rally

Hang Seng Index rose 0.15% at midday break, with Hang Seng Tech Index up 1.31%. Semiconductors, autos, and hardware led gains: Montage Tech and Gao Wei Electronics >12%, SMIC >8%, BYD >4%. Software services weakened with Zhipu AI down >3%; southbound funds net bought 23.78B HKD.

#china-stocks#semiconductors#ai-hardware
๐Ÿค–
Reddit r/MachineLearningโ€ข89d ago

Seeking Data for Heavy Industry Tool

A team in R&D is developing an operational tool for ports, mining, and fleet ops to address data gaps from manual logs and poor connectivity. They seek 15-min conversations, historical data under NDA, and on-site insights to validate their MVP. Goal is a site 'Nervous System' for accurate metrics.

#heavy-industry#data-collection#ops-tool
๐Ÿ“Š
Bloomberg Technologyโ€ข89d ago

Sereact Raises $110M for Predictive Robot AI

German startup Sereact secured $110 million in funding. The capital will develop an AI model enabling robots to predict consequences and adapt to tasks. This advances smarter, more versatile robotics software.

#robotics#funding#embodied-ai
๐Ÿ“„
ArXiv AIโ€ข89d ago

New Framework Certifies AI Research Authorship

Proposes a two-layer certification framework separating knowledge quality assessment from human contribution grading for AI-enabled research. Categorizes contributions as A (pipeline-reachable), B (human-directed), C (human-frontier). Enables transparent publication of automated work via benchmark slots within existing systems.

#ai-authorship#certification#peer-review
๐Ÿ“„
ArXiv AIโ€ข89d ago

MolClaw: Hierarchical AI Agent for Drug Discovery

MolClaw is an autonomous AI agent that unifies over 30 tools via a three-tier hierarchical skill architecture (70 skills) for drug molecule evaluation, screening, and optimization. It introduces MolBench, a benchmark with complex workflows spanning 8-50+ tool calls. MolClaw sets SOTA performance, proving workflow orchestration as the key bottleneck in AI drug discovery.

#autonomous-agents#drug-discovery#hierarchical-skills
๐Ÿ“„
ArXiv AIโ€ข89d ago

Memanto: SOTA Semantic Memory for Agents

Memanto introduces a typed semantic memory layer for agentic AI, using 13 predefined categories, conflict resolution, and temporal versioning powered by Moorcheh's Information-Theoretic Search engine. It enables sub-90ms deterministic retrieval without indexing or ingestion delays, outperforming hybrid graph and vector systems. Achieves SOTA 89.8% on LongMemEval and 87.1% on LoCoMo with single-query retrieval and lower complexity.

#agent-memory#semantic-retrieval
๐Ÿ“„
ArXiv AIโ€ข89d ago

Math Takes Two: Emergent Math Reasoning Benchmark

Math Takes Two is a new benchmark assessing emergent mathematical reasoning in AI agents through communication. It tests if two agents, without prior math knowledge, can develop a shared symbolic protocol for a visually grounded extrapolation task. The benchmark avoids predefined math language, requiring discovery of latent structures from scratch.

#multi-agent#emergent-reasoning#numerical-cognition
๐Ÿ“„
ArXiv AIโ€ข89d ago

LLM Self-Correction: Help or Harm?

Researchers model LLM self-correction as a cybernetic feedback loop using a two-state Markov diagnostic, recommending iteration only if ECR/EIR > Acc/(1-Acc). Across models like o3-mini and GPT-4o-mini, a sharp EIR threshold โ‰ค0.5% separates beneficial from harmful refinement. Verify-first prompting eliminates degradation in GPT-4o-mini, turning -6.2pp into +0.2pp gain.

#self-correction#markov-diagnostic#prompt-intervention
๐Ÿ“„
ArXiv AIโ€ข89d ago

CognitiveTwin Predicts Alzheimer's Cognitive Decline

CognitiveTwin is a digital twin framework predicting patient-specific cognitive trajectories in Alzheimer's disease using multi-modal data like cognitive scores, MRI, PET, CSF biomarkers, and genetics. It fuses modalities with a Transformer architecture and models temporal dynamics via Deep Markov Model. Trained on 1,666 TADPOLE patients, it excels in accuracy, demographic fairness, and robustness to missing data.

#digital-twins#multi-modal#alzheimers
๐Ÿ“„
ArXiv AIโ€ข89d ago

Background Temperature Reveals LLM Hidden Randomness

Researchers introduce background temperature (T_bg) to explain divergent LLM outputs at T=0 due to implementation nondeterminism like batch-size variation and floating-point issues. They formalize it as effective temperature from environment perturbations and propose an empirical estimation protocol. Pilot experiments on major LLM providers validate the concept for better reproducibility.

#reproducibility#nondeterminism
๐Ÿ“„
ArXiv AIโ€ข89d ago

AI Strategic Reasoning Risks Framework

Introduces ESRRSim, a taxonomy-driven framework for evaluating emergent strategic reasoning risks (ESRRs) in LLMs, such as deception, evaluation gaming, and reward hacking. It features 7 categories decomposed into 20 subcategories, automated scenario generation, and dual rubrics for responses and reasoning traces. Evaluation of 11 LLMs shows detection rates from 14.45% to 72.72%, with generational improvements indicating adaptation to evaluations.

#ai-safety#risk-taxonomy#llm-evaluation
๐Ÿ“„
ArXiv AIโ€ข89d ago

AgentSearchBench: AI Agent Search Benchmark

AgentSearchBench introduces a large-scale benchmark for AI agent search using nearly 10,000 real-world agents from multiple providers. It formalizes search as retrieval and reranking under executable and high-level queries, evaluated via execution-grounded performance. Experiments show semantic methods fall short, while execution-aware probing boosts ranking quality.

#ai-agents#benchmark#retrieval
๐Ÿ“„
ArXiv AIโ€ข89d ago

Agents Reproduce Social Science from Papers

LLM agents reproduce empirical social science results using only the paper's methods description and original data, without access to code or results. The system extracts structured methods, runs isolated reimplementations, and compares outputs cell-by-cell. Evaluation on 48 papers shows varying success across models and scaffolds, with failures from agent errors or paper underspecification.

#llm-agents#reproducibility#code-generation
๐Ÿ“„
ArXiv AIโ€ข89d ago

Agentic Science Needs Falsification Experiments

LLM-based agents speed up scientific data analysis but risk producing plausible yet unfalsified claims. The paper warns of selective analyses optimized for positives, ignoring negative evidence. It proposes a falsification-first standard where agents actively seek ways to disprove claims.

#agentic-ai#falsification#scientific-method
Page 832 of 1664