All Updates

Page 893 of 1661

April 21, 2026

🇨🇳
cnBeta (Full RSS)93d ago

Microsoft: Most Windows 11 Users Skip Third-Party AV

Microsoft officially states that most Windows 11 users with regular updates, default security settings, and normal habits can rely solely on built-in Windows Defender for adequate protection. This addresses long-standing debates on the need for third-party antivirus software. The feature is presented via Windows Security Center.

#endpoint-security#av-alternatives#os-protection
🇨🇳
cnBeta (Full RSS)93d ago

Xbox Helix: Powerful PC-Console Hybrid Revealed

Microsoft's next-gen Xbox Project Helix fuses PC and console gaming capabilities. It runs both PC and Xbox games, akin to Valve's upcoming Steam Machine. Exact form factor details remain undisclosed.

#gaming-hardware#pc-hybrid#console-strategy
🇨🇳
cnBeta (Full RSS)93d ago

AI HBM Demand Surges, 60% Short to 2027

Generative AI is fueling explosive demand for high-bandwidth memory (HBM). Chip vendors are ramping production aggressively, yet global capacity will only meet about 60% of needs by 2027. AI dominance in memory markets persists.

#supply-chain#chip-production#memory-shortage
🏠
IT之家93d ago

MIIT Pushes 5G+Industrial Internet Upgrade

China's MIIT is accelerating scaled development of 5G+Industrial Internet, completing 'hundreds-thousands-tens-of-thousands' 5G factory goals with over 25,000 projects and 1,260 factories built. Metrics show 20.5% quality improvement, 18.4% cost reduction, 24.7% capacity increase. Plans include new standards, guidelines, and industrial internet-AI fusion actions targeting 10,000 factories by 2027.

#china-policy#manufacturing-ai#5g-factories
⚛️
量子位93d ago

Liu Zhiyi Enters 2025 Forbes China Sci-Tech List

Liu Zhiyi, dean of Tianli Qiming AI Research Institute, has been selected for the 2025 Forbes China Sci-Tech Innovators list. The recognition highlights his leadership in blending education with AI to pioneer new paradigms in the intelligent era.

#ai-education#china-ai#forbes-list
🗾
ITmedia AI+ (日本)93d ago

Siemens Launches Autonomous Engineering AI Agent

Siemens announced Eigen Engineering Agent, a new AI product showcased at Hannover Messe 2026. It autonomously executes engineering tasks end-to-end—planning, execution, and verification—within actual engineering systems. This advances industrial AI beyond advisory functions into full autonomy.

#industrial-ai#agentic-ai
🦙
Reddit r/LocalLLaMA93d ago

Ditching Opus 4.7 for Kimi 2.6 Speed

A Max subscriber switches team from Anthropic's Opus 4.7 to Kimi 2.6 due to laziness, high costs, and errors. Kimi 2.6 offers faster, more reliable performance despite smaller context. User bought yearly sub and submitted PR for Forge integration.

#model-comparison#coding-llm#cli-tools
🇬🇧
The Guardian Technology93d ago

Anthropic Withholds Risky Mythos AI Model

Anthropic created Mythos Preview, a powerful model excelling at software vulnerabilities, but withheld public release for safety. Experts question its capabilities amid PR skepticism. Discussion explores strategy and regulation impacts.

#ai-safety#model-withhold#security
📄
ArXiv AI93d ago

WORC Optimizes Weak Links in Multi-Agent AI

WORC is a framework addressing reasoning instability in LLM multi-agent systems by identifying and reinforcing weak agents. It uses a two-stage process: meta-learning for zero-shot weak agent detection via task features and swarm intelligence, followed by uncertainty-driven extra reasoning budgets for weak links. Experiments show 82.2% average accuracy on benchmarks with improved stability and generalization.

#multi-agent#weak-link#llm-stability
📄
ArXiv AI93d ago

Unifying Memory, Skills, Rules in LLM Agents

Proposes the Experience Compression Spectrum framework, positioning memory, skills, and rules along a compression axis (5-20x for memory, 50-500x for skills, 1000x+ for rules) to manage LLM agent experience efficiently. Citation analysis of 1,136 references shows <1% cross-community citations and fixed compression in 20+ systems, missing adaptive cross-level compression. Highlights gaps in evaluation, transferability, and knowledge lifecycle management.

#agent-memory#skill-discovery
📄
ArXiv AI93d ago

SocialGrid Benchmark for Multi-Agent Social Reasoning

SocialGrid introduces an embodied multi-agent benchmark inspired by Among Us to evaluate LLM agents on planning, task execution, and social reasoning. Top models like GPT-OSS-120B achieve under 60% task completion, struggling with navigation and deception detection. It features a Planning Oracle to isolate social skills and provides failure analysis plus a competitive leaderboard.

#multi-agent#embodied-ai#benchmark
📄
ArXiv AI93d ago

SCF Fixes Multi-Agent LLM Conflicts

Multi-agent LLM systems fail 41-86.7% due to Semantic Intent Divergence. SCF middleware detects conflicts in real-time and resolves them via process-aware components, achieving 100% workflow completion across AutoGen, CrewAI, and LangGraph. It provides governance audit trails and is protocol-agnostic.

#multi-agent#conflict-resolution#semantic-consensus
📄
ArXiv AI93d ago

Rigorous XAI via Feature Attribution

This new arXiv paper critiques non-symbolic XAI methods like Shapley values and SHAP for lacking rigor, potentially misleading in high-stakes ML. It overviews ongoing efforts to adopt rigorous symbolic methods for reliable feature importance attribution.

#xai#explainability#symbolic-methods
📄
ArXiv AI93d ago

ReactBench: MLLM Topological Reasoning Benchmark

ReactBench introduces a benchmark for testing topological reasoning in Multimodal Large Language Models using chemical reaction diagrams. It features 1,618 expert-annotated QA pairs across four hierarchical task dimensions, from linear to cyclic structures. Evaluations on 17 MLLMs reveal a 30% performance gap in holistic structural reasoning, confirmed as a reasoning bottleneck via ablations.

#benchmark#chemical-diagrams#structural-reasoning
📄
ArXiv AI93d ago

MEDLEY-BENCH: Scale Boosts Evaluation, Not Control

MEDLEY-BENCH introduces a benchmark for AI metacognition, separating reasoning, private self-revision, and social revision under inter-model disagreement. It evaluates 35 models from 12 families across five domains, revealing scale improves evaluation but not control abilities. Smaller models often match or outperform larger ones, challenging scale as the sole path to metacognitive competence.

#metacognition#ai-benchmark#scale-laws
📄
ArXiv AI93d ago

MARCH: Multi-Agent CT Report Generator

MARCH is a multi-agent framework that emulates radiology department hierarchies to generate accurate 3D CT reports, addressing hallucinations in VLMs. It features a Resident Agent for initial drafting, Fellow Agents for revisions, and an Attending Agent for consensus. It outperforms SOTA on RadGenome-ChestCT in clinical fidelity and linguistic accuracy.

#multi-agent#radiology#medical-ai
📄
ArXiv AI93d ago

LLM Competency Questions Cross-Domain Study

This arXiv paper analyzes LLM-generated Competency Questions (CQs) for ontology engineering using quantitative measures like readability, relevance, and structural complexity. It compares open models (KimiK2-1T, Llama3.1-8B, Llama3.2-3B) and closed models (Gemini 2.5 Pro, GPT 4.1) across use cases. Results reveal distinct generation profiles shaped by use cases.

#competency-questions#ontology-engineering#llm-evaluation
📄
ArXiv AI93d ago

Graph-LLM-Agent Survey: Reasoning & Retrieval

This arXiv survey categorizes graph-LLM integrations by purpose (reasoning, retrieval, generation, recommendation), graph types (knowledge, scene, interaction, causal, dependency), and strategies (prompting, augmentation, training, agents). It maps methods across domains like cybersecurity, healthcare, finance, robotics. Offers practical guide for selecting best-fit approaches based on tasks and data.

#graphs#agent#rag
📄
ArXiv AI93d ago

DAP: Open-Source Hard Mode ATP Framework

DAP is an agentic framework using LLMs with self-reflection to discover theorem answers before formal proving in Lean 4's Hard Mode. It releases expert-reannotated MiniF2F-Hard and FIMO-Hard benchmarks. DAP achieves SOTA, solving 10 CombiBench problems and first to prove 36 PutnamBench theorems formally.

#theorem-proving#agentic-framework#benchmarks
📄
ArXiv AI93d ago

AAGMM: Governance Model for AI Agent Sprawl

This arXiv paper introduces the Agentic AI Governance Maturity Model (AAGMM), a five-level framework across 12 domains to manage AI agent sprawl in enterprises, grounded in NIST AI RMF and ISO/IEC 42001. It proposes a taxonomy of sprawl patterns like shadow agents and permission creep, linked to cost models. Validated through 750 simulations, it shows Level 4-5 maturity yields 94.3% lower sprawl, 96.4% fewer risks, and 32.6% higher task completion.

#ai-governance#agent-sprawl#maturity-model
Page 893 of 1661