All Updates
Page 880 of 1662
April 22, 2026
Neuro-Symbolic NARS Reasoning Benchmark
Introduces NARS-Reasoning-v0.1 benchmark pairing natural language with FOL, Narsese, and labels for reasoning. Develops FOL-to-Narsese pipeline validated in OpenNARS. Releases Phi-2 LoRA adapter for three-label classification.
LLMs Execute Science but Skip Reasoning
LLM-based scientific agents run workflows and inquiries but ignore evidence in 68% of traces and rarely revise beliefs via refutation. Base models dominate performance (41.4% variance) over scaffolds (1.5%). True scientific reasoning needs direct training, as outcomes alone can't validate results.
Human-Guided Harm Recovery for Agents
Researchers formalize harm recovery for LM agents acting on computers, steering from harmful to safe states per human preferences. A user study yields a dataset of 1,150 judgments and a rubric, powering a reward model that re-ranks recovery plans. BackBench benchmark with 50 tasks shows improved recovery over baselines.
GROVE Visualizes LM Output Distributions
GROVE is an interactive visualization tool that depicts multiple language model generations as overlapping paths in a text graph, exposing shared structures, branches, and clusters. Informed by a study with 13 LM researchers, it addresses limitations of single-output interactions for prompt iteration. User studies confirm it improves structural judgments like diversity assessment.
Error-Free Training on MedMNIST Datasets
Researchers introduce Artificial Special Intelligence, a method for error-free training of ML classification models by avoiding repeated mistakes. Applied to 18 MedMNIST biomedical image datasets, it achieves perfect accuracy on 15, with three failing due to double-labeling issues. The work is detailed in arXiv preprint 2604.18916v1.
AutomationBench: AI Workflow Benchmark Launch
AutomationBench introduces a benchmark for AI agents tackling cross-application workflows via REST APIs, addressing gaps in coordination, API discovery, and policy adherence. Tasks draw from real Zapier patterns across Sales, Marketing, Operations, Support, Finance, and HR. Frontier models score below 10% on programmatic end-state grading amid misleading data.
ARES Fixes RLHF Dual Safety Flaws
ARES introduces a framework to discover and fix systemic weaknesses in RLHF where both LLMs and Reward Models fail together. It uses a Safety Mentor to generate adversarial prompts and responses targeting both components. A two-stage repair process fine-tunes the RM first, then optimizes the core model, boosting safety on benchmarks without capability loss.
AI + Lean 4 Verifies Patent Claims
Introduces first formally verified framework for patent analysis via hybrid AI + Lean 4 pipeline. Machine-checks DAG-coverage core and formalizes IP use cases like freedom-to-operate and doctrine-of-equivalents. Certifies computations post-ML scores, bridging AI with interactive theorem proving.
Adversarial Environments Fool Agentic AI
Tool-integrated AI agents face a Trust Gap, evaluated for capability but not skepticism against lying tools. Researchers formalize Adversarial Environmental Injection (AEI) and release POTEMKIN harness for robustness testing. Tests on 11,000+ runs reveal trade-offs between resisting Illusion (epistemic) and Maze (navigational) attacks.
Autonomous Driving Battle Heats Up Pre-2026 Beijing Auto Show
Ahead of the 2026 Beijing Auto Show, autonomous driving competition intensifies as carmakers vie for dominance. The event is positioned as a key battleground, or 'Bright Summit.' New developments are being foreshadowed.
Top AI Disease Mgmt Firm Races to IPO
The market-leading company in AI full disease course management is pushing for an IPO. It has secured investments from funds backed by executives from Baidu Group, Alibaba, and Tencent. This underscores strong big-tech interest in AI healthcare.
China Industrial AI Deployment Stuck
The article explores bottlenecks preventing most AI from entering production lines in China. It questions why deployable industrial AI remains scarce despite potential.
Two Chip Giants Double Down on China
Two leading chip giants are making significant investments in China. Emphasizes a strategy of pursuing opportunities on both domestic and international fronts.
10 Tips to Master OpenAI Images 2.0
OpenAI has launched Images 2.0, hailed as the new king of image generation. The article offers hands-on testing with 10 practical tips to fully utilize it. It warns that the real risks go beyond improved drawing capabilities.
Benchmarks Needed for DeepSeek V3.2 Quants
Developer seeks benchmarks to measure quality loss from quantization on DeepSeek V3.2 for a runtime quantization product. Focus is comparing quantized vs full-precision performance.
Qwen3.6-35B-A3B Local Setup on M2 Mac
User shares working config for running Qwen3.6-35B-A3B via llama.cpp on MacBook Pro M2 Max with pi coding agent. Setup uses UD-Q5_K_XL quant (~19GB), 128k context, and OpenAI-compatible API. Includes full llama-server command and models.json for pi agent integration.
Apple Hunts for AI Genius
Apple is in need of an exceptional AI talent. Explores how Apple can 'Think Different' in the AI landscape.
ChatGPT Images 2.0: AI Thinks Before Drawing
OpenAI announced ChatGPT Images 2.0, an advanced image generation model where AI reasons step-by-step before creating visuals. It significantly improves accuracy for Japanese text rendering compared to previous versions.
iQIYI Dares Tech Push Amid Backlash
iQIYI founder Gong Yu pushes a technical scheme despite public outcry. He underestimated emotional public response. The strategy has ignited widespread debate.
Meta Staff Rebel Against AI Surveillance PCs
Meta is installing surveillance software on employees' work computers to capture keystrokes for AI training. Staff are unhappy, highlighting the irony given Meta's user-tracking business model. The move is reportedly driven by Zuckerberg to fuel AI development.