SetupAI

Daily AI briefing

This WeekToolsUpdatesSearch繁
繁

Full archive

Every story we have kept, newest first.

Looking for the daily editions? → Past editions

Page 1360 of 1372

February 13, 2026

CausalAgent: Conversational Causal Inference
Applications

CausalAgent: Conversational Causal Inference

CausalAgent is a multi-agent system automating end-to-end causal inference via natural language. Integrates MAS, RAG, and MCP for data cleaning to report generation.

ArXiv AI · 214d ago

C-JEPA Learns World Models via Object Masking
Image

C-JEPA Learns World Models via Object Masking

C-JEPA extends masked joint embedding prediction to object-centric representations with object-level masking, inducing latent interventions for interaction reasoning. It boosts counterfactual VQA by 20% and enables efficient agent planning using 1% of latent features.

ArXiv AI · 214d ago

Bi-Level Optimization for Multimodal LLM Judges
Models

Bi-Level Optimization for Multimodal LLM Judges

Introduces BLPO to optimize prompts for multimodal LLM-as-a-judge evaluating AI images. Overcomes context limits by converting images to text representations.

ArXiv AI · 214d ago

BHI Framework Audits LLM Benchmarks
Models

BHI Framework Audits LLM Benchmarks

Introduces Benchmark Health Index (BHI), a data-driven framework to audit LLM benchmarks amid reliability issues like score inflation. Evaluates along three axes: Capability Discrimination, Anti-Saturation, and Impact.

ArXiv AI · 214d ago

Benchmarking LLM Agents Under Noise
Models

Benchmarking LLM Agents Under Noise

AgentNoiseBench evaluates tool-using LLM agents' robustness in noisy real-world environments. Categorizes noise into user-noise and tool-noise; injects controllable perturbations into benchmarks.

ArXiv AI · 214d ago

Benchmark for LLM Replication in Sciences
Models

Benchmark for LLM Replication in Sciences

ReplicatorBench tests LLM agents on replicating social/behavioral science claims end-to-end. Covers extraction, experiments, and interpretation with replicable/non-replicable cases.

ArXiv AI · 214d ago

Behavioral Optimization for Proactive Agents
Models

Behavioral Optimization for Proactive Agents

BAO uses agentic RL to train proactive LLM agents balancing performance and user engagement. Combines behavior enhancement with regularization to align with user expectations.

ArXiv AI · 214d ago

AT-RL Reinforces MLLM Anchors for Reasoning
Models

AT-RL Reinforces MLLM Anchors for Reasoning

AT-RL selectively reinforces high-connectivity cross-modal anchor tokens (15% of total) in MLLM RLVR via attention graph clustering. 32B model hits 80.2% on MathVista, beating 72B baseline with 1.2% overhead.

ArXiv AI · 214d ago

ARC Learns Dynamic Agent Configurations
Models

ARC Learns Dynamic Agent Configurations

ARC introduces a reinforcement learning policy to dynamically configure LLM-based agent systems per query, selecting optimal workflows, tools, and prompts. It outperforms fixed templates on reasoning and tool-augmented QA benchmarks.

ArXiv AI · 214d ago

AIR Boosts LLM Agent Safety
Models

AIR Boosts LLM Agent Safety

AIR is the first incident response framework for LLM agents, focusing on detecting, containing, recovering from, and eradicating incidents post-occurrence. It integrates a domain-specific language into the agent's execution loop for autonomous management.

ArXiv AI · 214d ago

AgentLeak: Multi-Agent Privacy Leak Benchmark
Models

AgentLeak: Multi-Agent Privacy Leak Benchmark

AgentLeak introduces the first full-stack benchmark for privacy leakage in multi-agent LLM systems, covering internal channels like inter-agent messages. It spans 1,000 scenarios across healthcare, finance, legal, and corporate domains.

ArXiv AI · 214d ago

Gemini 3 Deep Think Dominates Programming
Coding

Gemini 3 Deep Think Dominates Programming

Gemini 3 Deep Think receives a major upgrade, achieving state-of-the-art results across domains, especially programming. Only 7 people globally outperform it.

cnBeta (Full RSS) · 214d ago

Gemini 3 Deep Think Dominates Coding
Coding

Gemini 3 Deep Think Dominates Coding

Gemini 3 Deep Think upgrade achieves SOTA across domains, especially programming where only 7 people worldwide outperform it. This Google VP side project marks a new era in AI reasoning.

cnBeta (Full RSS) · 214d ago

Ex-Researcher Warns on ChatGPT Ads
Business

Ex-Researcher Warns on ChatGPT Ads

Former OpenAI researcher Zoë Hitzig warns ads in ChatGPT risk user manipulation like Facebook. She left after ad testing amid privacy concerns from user-shared intimate thoughts.

cnBeta (Full RSS) · 214d ago

Ex-Researcher Warns ChatGPT Ads Risks
Business

Ex-Researcher Warns ChatGPT Ads Risks

Former OpenAI researcher Zoë Hitzig quit after testing ChatGPT ads, warning of manipulation risks from users' private data. She compares it to Facebook's pitfalls.

cnBeta (Full RSS) · 214d ago

OpenAI Adds Ads to ChatGPT
Business

OpenAI Adds Ads to ChatGPT

OpenAI is launching ads on ChatGPT this week amid billions in funding needs. CEO Sam Altman previously opposed ads, calling them a last resort that could erode user trust.

cnBeta (Full RSS) · 214d ago

Gemini 3 Deep Think: Sketch to 3D
Image

Gemini 3 Deep Think: Sketch to 3D

Google unveiled a major upgrade to Gemini 3 Deep Think, a reasoning model for science, research, and engineering. Google AI Ultra subscribers can access it now in the Gemini App.

cnBeta (Full RSS) · 214d ago

Gemini 3 Deep Think: Sketch-to-3D Upgrade
Image

Gemini 3 Deep Think: Sketch-to-3D Upgrade

Google announced a major upgrade to Gemini 3 Deep Think, a reasoning model for science, research, and engineering. Google AI Ultra subscribers can access it via the Gemini App.

cnBeta (Full RSS) · 214d ago

Anthropic Targets OpenAI in Super Bowl Ads
Business

Anthropic Targets OpenAI in Super Bowl Ads

AI companies with deep user privacy access are rushing to monetize via ads amid lax regulation. Anthropic's Super Bowl ads satirized OpenAI's vulnerabilities without naming them.

cnBeta (Full RSS) · 214d ago

AI Safety Leader Quits for Poetry
Research

AI Safety Leader Quits for Poetry

An AI safety leader warns of global peril and resigns to study poetry. This coincides with an OpenAI researcher quitting over ChatGPT ad testing plans.

BBC Technology · 214d ago

AI Safety Chief Quits, Cites Global Peril
Research

AI Safety Chief Quits, Cites Global Peril

A prominent AI safety leader resigned, warning the world is in peril, to study poetry. This follows an OpenAI researcher's exit over plans to test ChatGPT ads.

BBC Technology · 214d ago

Samsung First Ships HBM4 Memory
Infrastructure

Samsung First Ships HBM4 Memory

Samsung claims first to ship HBM4 memory, a day after Micron's announcement. HBM4 provides faster, denser RAM for next-gen AI hardware.

The Register - AI/ML · 214d ago

Agent Frameworks Essential Despite LLM Gains
Infrastructure

Agent Frameworks Essential Despite LLM Gains

Discusses if agent frameworks remain necessary as LLMs improve. Argues building approaches evolve but agents are fundamentally systems around models.

LangChain Blog · 214d ago

Agent Frameworks Remain Vital Amid LLM Advances
Infrastructure

Agent Frameworks Remain Vital Amid LLM Advances

Explores whether agent frameworks are still necessary as LLMs improve. Notes that optimal agent-building approaches evolve with model performance.

LangChain Blog · 214d ago

MiniMax Drops Affordable M2 AI Model
Models

MiniMax Drops Affordable M2 AI Model

MiniMax released an updated M2 large language model for real-world productivity. The cheap AI follows rivals' launches in China's intense AI race.

SCMP Technology · 214d ago

AI Advances via Compute, Not Smarts
Research

AI Advances via Compute, Not Smarts

MIT report shows frontier models like OpenAI's GPT rely on more computing power rather than smarter algorithms. This scaling approach drives progress but hikes costs.

ZDNet AI · 214d ago

Xiaomi Open-Sources Robot VLA Model
Open Source

Xiaomi Open-Sources Robot VLA Model

Xiaomi open-sources its first-generation VLA large model for robotics. Part of morning tech news roundup alongside OpenAI updates.

Ifanr (爱范儿) · 214d ago

AI Coding Flaws Hack BBC Reporter
Coding

AI Coding Flaws Hack BBC Reporter

Security flaws in an AI coding platform allowed a BBC reporter to be hacked. Vibe-coding tools enable non-coders to build apps using AI.

BBC Technology · 214d ago

Cloudflare Serves Markdown to AI
Infrastructure

Cloudflare Serves Markdown to AI

Cloudflare optimizes websites for AI agents by converting HTML to Markdown. Shifts from blocking bots to attracting them with faster content.

The Register - AI/ML · 214d ago

AI Supercharges Call Center Agents
Applications

AI Supercharges Call Center Agents

UJET CEO claims AI enhances call agents into 'superheroes' without job loss. Improves software to resolve issues without multi-system navigation.

The Register - AI/ML · 214d ago

11359136013611372
Page 1360 of 1372
Back to home
SetupAI

A bilingual daily AI briefing — ten stories a day, each with a deep insight.

Takedown / opt-out: copyright@setupai.uk

© 2026 SetupAI

This WeekToolsUpdatesAboutPrivacyTermsRSS