Latest Multimodal AI News & Updates

Models that see, hear and speak are replacing text-only systems. Coverage of vision-language models, native audio and video understanding.

303 articles

🐯
虎嗅75d ago

Google launches Gemini 3.5 Flash and AI agent ecosystem

Google unveiled Gemini 3.5 Flash, the Omni multi-modal model, and the Spark AI assistant at I/O 2026. The company is aggressively integrating these tools across its entire product suite to dominate the AI agent market.

🐯
虎嗅76d ago

Google I/O: AI Integration Across the Entire Ecosystem

Google's I/O conference focused on integrating AI across its entire product suite, introducing the Gemini Omni world model and the cost-effective Gemini 3.5 Flash to drive enterprise and consumer adoption.

🧬
DeepMind Blog78d ago

Introducing Gemini Omni

Google officially introduces Gemini Omni, a new multimodal model designed for real-time, low-latency interaction. This model enhances the capabilities of the Gemini ecosystem by providing seamless cross-modal processing.

🦙
Reddit r/LocalLLaMA123d ago

Gemma 4 Released: Multimodal Open Models

Google DeepMind launched Gemma 4, open-weights multimodal models in E2B, E4B, 26B A4B, and 31B sizes. They handle text, images, video, and audio, with up to 256K context and strong reasoning, coding, agentic features. Dense and MoE variants optimized for on-device deployment.

🏠
IT之家21d ago

Xihu University and Alibaba DAMO Academy Launch Guiyuan AI

Xihu University and Alibaba DAMO Academy have developed 'Guiyuan', an AI model designed to predict stem cell fate. By analyzing millions of molecular combinations, the model enables precise control over cell reprogramming, significantly accelerating biological research.

🏠
IT之家22d ago

SenseNova-Vision: Unified Open-Source Visual Foundation Model

SenseTime has released SenseNova-Vision, a unified visual foundation model that integrates tasks like detection, segmentation, and 3D reconstruction. It is fully open-source, including a 50 million sample visual instruction corpus.

📄
ArXiv AI25d ago

Infinity-Parser2: New SOTA Multimodal Document Parsing Model

Infinity-Parser2 is a new multimodal model featuring a 5-million-sample synthetic corpus and a multi-task reinforcement learning framework. It achieves state-of-the-art performance on document parsing benchmarks, outperforming models like DeepSeek-OCR-2.

🗾
ITmedia AI+ (日本)25d ago

Meta Releases Muse Spark 1.1 and New Low-Cost API

Meta has unveiled 'Muse Spark 1.1', a multimodal reasoning model with enhanced agentic capabilities for tool use and complex coding. The model is available via the new 'Meta Model API' at a price point lower than competing models.

📰
The Verge26d ago

Meta releases Muse Spark 1.1 for coding

Meta has launched Muse Spark 1.1, an upgraded AI model designed for advanced coding tasks. It features improved bug detection, support for multi-agent workflows, and native multimodal perception.

🔥
36氪27d ago

Tencent Hires OpenAI Researcher Yonglong Tian for VLM

Tencent has hired former OpenAI researcher Yonglong Tian to join its Large Language Model department. He will focus on the research and development of Vision-Language Models (VLM).

🔥
36氪32d ago

Shengshu Tech Launches Vidu S1 Real-time Interactive Model

Shengshu Tech has released Vidu S1, a new video generation model capable of real-time interaction, voice control, and high-definition video output.

⚛️
量子位38d ago

Global first: Edge-side streaming multimodal model released

A Hangzhou-based team has successfully deployed a streaming multimodal model to the edge, marking a significant milestone for CVPR 2026 trends.

📲
Digital Trends42d ago

Pixel-Level Photo Attacks Bypass AI Chatbot Safety Rules

Researchers at Florida International University discovered that invisible pixel-level modifications in images can force AI chatbots to bypass their safety guardrails and generate restricted content.

🧬
DeepMind Blog56d ago

Introducing Gemma 4 12B: A Unified Multimodal Model

Google DeepMind has released Gemma 4 12B, a new encoder-free multimodal model. This release expands the Gemma family with enhanced capabilities for processing diverse data types.

🇨🇳
cnBeta (Full RSS)61d ago

Google Releases Open-Source Gemma 4 12B Multimodal Model

Google has released the Gemma 4 12B multimodal model, designed to run locally on consumer hardware. It offers performance comparable to the 26B version while requiring only 16GB of memory.

🗾
ITmedia AI+ (日本)61d ago

Google Releases Gemma 4 12B Multimodal Model for Local PCs

Google has launched 'Gemma 4 12B', an open multimodal model designed for efficiency. It features an encoder-free architecture that allows it to run on laptops with 16GB of RAM.

⚛️
量子位64d ago

World's first omni-modal API now free and open

A top-tier AI lab has released the world's first omni-modal API, which is now available for free indefinitely. The API supports seamless processing of text, images, and video data.

💰
钛媒体76d ago

Google Gemini 3.5 and AI Industry Shifts

Google released Gemini 3.5 Flash and Omni models, while the industry sees major shifts with Anthropic hiring Andrej Karpathy and AMD expanding local AI infrastructure. AI competition is evolving from raw parameter counts to ecosystem building and industry penetration.

📱
Ifanr (爱范儿)76d ago

Google launches Gemini 3.5 and new AI agent products

Google has officially unveiled the Gemini 3.5 model alongside new AI agent capabilities and advanced video generation models. This release marks a significant shift in Google's product strategy toward autonomous AI agents.

🦙
Reddit r/LocalLLaMA81d ago

Intern-S2-Preview: Efficient 35B Scientific Multimodal Model

Intern-S2-Preview is a new 35B scientific foundation model that scales task difficulty and diversity. It features enhanced agentic capabilities and efficient reinforcement learning using CoT compression.

🦙
Reddit r/LocalLLaMA83d ago

Ovis2.6-80B-A3B: MoE multimodal model with active vision

AIDC-AI introduced Ovis2.6-80B-A3B, a multimodal LLM using a Mixture-of-Experts architecture. It features 'Think with Image' capabilities, allowing the model to actively manipulate visual inputs during reasoning.

🤗
Hugging Face Blog97d ago

NVIDIA Launches Nemotron 3 Nano Omni Multimodal Model

NVIDIA introduces Nemotron 3 Nano Omni, a compact long-context multimodal model for building intelligent agents handling documents, audio, and video. Hosted on Hugging Face, it advances efficient multimodal AI capabilities. This launch targets agent developers seeking lightweight yet powerful solutions.

🔥
36氪98d ago

SenseTime Open-Sources SenseNova U1 Models

SenseTime officially released and open-sourced the SenseNova U1 series, a native unified model for understanding and generation. Built on the in-house NEO-unify architecture from March, it integrates multimodal understanding, reasoning, and generation. The model treats language and vision as a unified composite for efficient synergy while preserving semantics and pixel fidelity.

💼
VentureBeat98d ago

Xiaomi MiMo-V2.5 Tops Agentic Claw Efficiency

Xiaomi launched open-source MiMo-V2.5 and V2.5-Pro models under MIT license, available on Hugging Face for commercial use. They lead in agentic 'claw' tasks like OpenClaw, with Pro achieving 63.8% success on ~70K tokens—40-60% fewer than Claude or GPT. Featuring 310B parameters and 1M-token context, they challenge closed-source frontiers.

🦙
Reddit r/LocalLLaMA103d ago

Qwen3.6-27B Uncensored Aggressive Launches

New fully uncensored Qwen3.6-27B Aggressive model released with optimized K_P GGUF quants. Achieves 0 refusals with no capability loss, supports multimodal inputs. Sensitive to prompt clarity; disable thinking via Jinja kwarg.

🦙
Reddit r/LocalLLaMA109d ago

Qwen3.6-35B Uncensored Aggressive Released

New uncensored aggressive variant of Qwen3.6-35B-A3B released with K_P quants and full multimodal support. Zero refusals, no capability loss, and optimized for local inference. Includes various quants, vision mmproj, and 262K context.

🌍
The Next Web (TNW)109d ago

Claude 4.7 Tops Coding Benchmarks

Anthropic launches Claude Opus 4.7, its top model with 64.3% on SWE-bench Pro (vs GPT-5.4's 57.7%), multi-agent coordination for long workflows, 3x image resolution, and 14% better agentic reasoning with fewer tool errors. Pricing: $5 input/$25 output per million tokens.

🦙
Reddit r/LocalLLaMA110d ago

Qwen3.6-35B-A3B Open-Source MoE Launched

Qwen3.6-35B-A3B is a new open-source sparse MoE model with 35B total parameters and 3B active params under Apache 2.0. It matches agentic coding of models 10x larger and offers strong multimodal perception with thinking/non-thinking modes. Available on HuggingFace, ModelScope, and Qwen Studio.

📊
Bloomberg Technology112d ago

Alibaba's Happy Horse Tops Video Gen Benchmarks

Alibaba stealthily released the Happy Horse AI video generator last week. It now ranks top on key video generation benchmarks. This positions Alibaba as the leader in AI video synthesis.

🦙
Reddit r/LocalLLaMA123d ago

Uncensored Gemma 4 E4B/E2B Multimodal Launch

New aggressive uncensored variants of Gemma 4 E4B (4B) and E2B (2B) released as fully multimodal models supporting text, image, video, and audio. Available in high-quality GGUF quants on Hugging Face, compatible with llama.cpp. Larger E31B and E26B models coming soon.