All Updates
Page 622 of 1670
May 14, 2026
State of Security 2026: Application Security Challenges
The integration of AI into software development workflows is creating new security vulnerabilities. Organizations must adapt their security strategies to keep pace with rapid AI-driven coding cycles.
REVELIO: Uncovering Interpretable Failure Modes in VLMs
REVELIO is a new framework designed to systematically identify interpretable failure modes in Vision-Language Models (VLMs). By combining diversity-aware beam search and Gaussian-process Thompson Sampling, it uncovers specific scenarios where models consistently fail, such as in autonomous driving or robotics.
PROMETHEUS: Automating Deep Causal Research with World Models
PROMETHEUS is a new framework that transforms unstructured research data, code, and literature into navigable 'causal atlases'. It enables researchers to synthesize local causal claims into a coherent Topos World Model, facilitating better evidence tracking and counterfactual testing.
Polynomial Complexity for First-Order Knowledge Base Progression
Researchers demonstrate that first-order progression for local-effect, normal, and acyclic action classes grows only polynomially. This finding ensures that knowledge base updates remain within decidable fragments, enhancing practical applicability in AI reasoning.
Multimodal HMMs for Persistent Emotional State Tracking
Researchers introduced a lightweight framework using sticky factorial HDP-HMMs to track emotional arcs in conversations. By analyzing multimodal valence-arousal data, the model identifies persistent emotional regimes more efficiently than LLM-based dialogue tracking.
MAVIC: Solving Instruction Conflicts in Multi-Agent Reinforcement Learning
MAVIC is a new reinforcement learning framework that resolves value estimation inconsistencies when external instructions interrupt macro-actions. By modifying the bootstrapping target, it enables agents to maintain high instruction compliance without sacrificing base task performance.
DisaBench: A Participatory Framework for Evaluating Disability Harms
DisaBench is a new participatory evaluation framework designed to identify subtle disability-related harms in large language models. It features a taxonomy of twelve harm categories co-created with people with disabilities and a dataset of 175 annotated prompts.
CLIPR: Learning Transferable Latent User Preferences for LLMs
CLIPR is a new framework that enables LLMs to infer and apply latent user preferences from minimal conversational interactions. It generates actionable, transferable natural language rules that improve decision-making alignment across various tasks and environments.
CHAL: A New Framework for Multi-Agent Dialectic Reasoning
CHAL is a novel multi-agent framework designed to improve LLM reasoning in defeasible domains by treating debate as structured belief optimization. It utilizes a Bayesian-inspired belief schema and configurable meta-cognitive value systems to ensure transparent and aligned reasoning.
BenchJack: Automating Red-Teaming to Expose AI Benchmark Flaws
BenchJack is an automated red-teaming system designed to identify reward-hacking exploits in AI agent benchmarks. By applying it to 10 major benchmarks, researchers discovered 219 flaws, demonstrating that many current evaluation metrics are easily gamed without actual task completion.
BEHAVE: Real-Time Modeling of Collective Human Dynamics
BEHAVE is a new AI framework that models groups of humans as complex dynamical systems. It uses kinematic micro-signals to map interaction graphs and forecast collective behavior in real-time.
Jiuzhang 4 Quantum Computer Achieves 10^54 Speedup
Researchers at the University of Science and Technology of China have unveiled Jiuzhang 4, a photonic quantum prototype. It demonstrates a quantum advantage 10^54 times faster than current supercomputers.
Microsoft Retires Copilot Mode in Edge Browser
Microsoft is phasing out the dedicated 'Copilot Mode' in the Edge browser. This change reflects the company's strategy to integrate Copilot features directly into the core browser experience rather than keeping them as a separate mode.
Takeda to cut 4,500 jobs by 2026
Takeda Pharmaceutical announced a major restructuring plan to cut approximately 4,500 jobs by fiscal year 2026. The move aims to streamline operations and reduce costs, representing about 10% of its total workforce.
Tencent Cloud's decline in the AI and MaaS era
Tencent Cloud has fallen to fifth place in China's public cloud market, struggling with late AI model deployment and limited GPU resources. The company is currently attempting a strategic pivot to regain competitiveness in the MaaS sector.
Continual Harness: Online adaptation for self-improving agents
The paper introduces 'Continual Harness,' a framework for model-harness co-learning that enables AI agents to refine their own tools and harnesses. This approach was successfully demonstrated by agents completing complex PokΓ©mon games through iterative self-refinement.
Retailers use weight loss challenges for user growth
Major retailers like JD Health and Yonghui Superstores are launching weight-loss challenges, using cash and food rewards to drive traffic and social media engagement.
Changan Mazda: Strategic turnaround under Wang Xiaoling
Changan Mazda is undergoing a critical strategic shift led by Wang Xiaoling to combat declining market share. The company is focusing on aggressive market recovery and product repositioning.
Korea urges Samsung-union talks on May 16
The National Labor Relations Commission of South Korea has requested Samsung Electronics and its labor union to resume negotiations on May 16. The government warns that potential strikes pose significant risks to economic growth and the semiconductor supply chain.
The Evolution of AI Translation and Human Agency
The article reflects on the rise of AI in translation, suggesting that we should embrace its capabilities while maintaining a natural, human-centric approach. It explores the philosophical shift in how we perceive translation tasks in an AI-augmented world.