All Updates
Page 816 of 1664
April 28, 2026
Vercel 2026 AI Accelerator Winners Revealed
Vercel's 2026 AI Accelerator concluded with Demo Day, where 39 global teams pitched AI apps in agents, dev tools, finance, security, and more. Each team received over $200K in credits; winners Rex (enterprise finance AI), Hacktron AI (security teammate), and Roots (real estate) took prizes and investments. Alumni raised $100M+ and joined Y Combinator.
Systematic LLM Debugging Approach
This paper introduces a systematic, model-agnostic approach for debugging large language models (LLMs), treating them as observable systems. It provides structured methods from issue detection to model refinement, unifying evaluation, interpretability, and error analysis. The methodology enables iterative diagnosis of weaknesses, prompt refinement, and data adaptation, even without standardized benchmarks.
Power Laws Boost AI Compositional Reasoning
Training AI models on power-law distributed natural language data outperforms uniform distributions across compositional reasoning tasks like state tracking and multi-step arithmetic. Power-law sampling creates beneficial asymmetry that flattens the loss landscape, enabling efficient learning of high-frequency skills first as a foundation for rare long-tail skills. Theoretical proofs show it requires significantly less data.
Polynomial Inverse for Preference Argumentation
New arXiv paper proves that finding preference relations yielding a desired labelling in preference-based argumentation frameworks (PAFs) is polynomial-time solvable for most common reductions under complete semantics. This inverse problem supports preference elicitation and explainability in AI argumentation. It analyzes four widely-used PAF-to-AAF reductions.
PhySE Framework for AR-LLM Social Attacks
PhySE tackles AR-LLM social engineering attack bottlenecks like cold-start profiling delays and static strategies. It uses VLM pre-training for rapid social profiles and an adaptive psychological LLM for dynamic tactics. Evaluated via IRB-approved study with 60 participants and 360 conversations.
PExA Hits 70.2% SOTA on Spider 2.0
PExA reformulates text-to-SQL as software test coverage using parallel atomic SQL test cases for semantic coverage. It generates the final SQL only after sufficient exploration, addressing LLM agent latency-performance trade-offs. Achieves new state-of-the-art 70.2% execution accuracy on Spider 2.0 benchmark.
Multi-Fidelity Digital Twins for Aircraft Fault Diagnosis
Proposes intelligent fault diagnosis for general aviation aircraft using multi-fidelity digital twins, FMEA fault injection, residual features, and LLM reports. Builds JSBSim-based twin generating 23-channel data for 19 fault types. Achieves 96.2% Macro-F1, emphasizing residual quality over classifiers.
Multi-Agent LLMs Automate Ontology Generation
Researchers propose a multi-agent LLM framework for generating ontologies from unstructured text, outperforming single-agent baselines on insurance contracts. It decomposes tasks into Domain Expert, Manager, Coder, and Quality Assurer roles. Results highlight planning-first, artifact-driven design for better structural quality and queryability.
Interpretable Wi-Fi HAR with Discrete Latents & LTL Rules
Proposes a pipeline compressing Wi-Fi CSI to discrete latents via categorical VAE, then applies causal discovery and LTL rule extraction for symbolic HAR classifier. Achieves competitive performance with full interpretability, no black-box components. Enables symbolic multi-antenna fusion without retraining.
Graphs Boost LLM Multi-Agent Reasoning
Researchers test explicit belief graphs for improving LLM performance in cooperative multi-agent reasoning using the Hanabi game across 3,000+ trials and four LLM families. Graphs prove essential when gating action selection, enabling 100% success on 2nd-order Theory of Mind for strong models versus 20% baseline. Key findings include 'Planner Defiance,' inter-agent conventions outperforming single-agent methods, and shallow graphs offering the best cost-benefit.
FormalScience: Scalable Science Autoformalisation in Lean
FormalScience is a human-in-the-loop agentic pipeline that enables domain experts to autoformalize scientific reasoning into syntactically correct Lean4 proofs at low cost. It introduces FormalPhysics, a dataset of 200 university-level physics problems with formal representations. The paper evaluates LLMs on autoformalisation, characterizes semantic drift, and releases an open-source codebase with interactive UI.
Decoupled HITL for Agent Autonomy
This arXiv paper introduces a decoupled Human-in-the-Loop (HITL) architecture that separates human oversight from agent workflows for scalability and reuse. It proposes a framework formalizing HITL integration across four dimensions: intervention conditions, role resolution, interaction semantics, and communication channels. The design enables protocol-level HITL for consistent governance in multi-agent environments.
Analytica: 15% LLM Reasoning Accuracy Boost
Analytica introduces Soft Propositional Reasoning (SPR) to stabilize LLM agents in complex analysis like forecasting. It decomposes tasks into subpropositions, grounds them with tool agents including a Jupyter Notebook agent, and synthesizes via linear models to cut bias and variance. Delivers 15.84% average accuracy gain, up to 71% with low 6% variance.
Welding Bots Raise Hyundai Funding, Ship Orders
昇視唯盛完成數千萬元A+輪融資,由韓國現代與微光創投領投,用於具身AI研發與產品迭代。第三代自主移動焊接機器人已於2025年推出,效率提升30%,並簽下船廠數千萬元訂單。2026年营收預計破數億。
User Abandons Local LLMs for Coding
A developer tested top local LLMs like Qwen 27B and Gemma 31B for coding and OS tasks but found them inferior to Claude. Key issues include poor decision-making, unreliable tool calls during Docker builds, and performance lags from prompt cache failures. Concludes productivity loss outweighs local advantages.
China Fiber Optics Boom from AI Demand
Global optical fiber demand surges due to AI data centers requiring 5-10x more fiber, causing prices to skyrocket—Chinese G.657.A2 up 8x to 240 RMB/km. Chinese firms hold 60% capacity with orders booked to 2027Q1. Production lags from 18-24 month preform rod cycles.
Traditional Carmakers Falter in Huawei EV Deals
BAIC BluePark revenue doubles but posts 4.5B RMB loss amid high R&D/channel costs for Enjoy/ArcFox despite Huawei tech boost. JAC's Zunjie hits sales milestones but drags overall with idle capacity and Huawei marketing fees. OEMs act as low-margin manufacturers under Huawei control.
IPO Rules Ease Exits for Loss-Making AI Tech
VC panel notes secondary market recovery favors high-tech like semis/GPUs despite losses, with IPOs now open to quantum/BCI firms. Investors shift to policy-backed 'new quality forces'; patience key for tech moats. Valuations rebound as exits smooth.
DeepSeek V4 Report Reveals R&D Departures
DeepSeek V4's 58-page technical report lists nearly 300 authors, with 10 marked as departed. At least five core R&D members left since late 2025, affecting base models, reasoning, OCR, and multimodal areas. This has sparked attention on the company's talent stability.
Humanoid Robots Trial as Airport Baggage Handlers
Japan Airlines will trial humanoid robots at Tokyo's Haneda airport from May to address labor shortages amid surging tourism. The robots will assist overburdened baggage handlers but require regular recharging.