Latest Multimodal AI News & Updates
Models that see, hear and speak are replacing text-only systems. Coverage of vision-language models, native audio and video understanding.
303 articles
Google launches Gemini 3.5 Flash and AI agent ecosystem
Google unveiled Gemini 3.5 Flash, the Omni multi-modal model, and the Spark AI assistant at I/O 2026. The company is aggressively integrating these tools across its entire product suite to dominate the AI agent market.
Google I/O: AI Integration Across the Entire Ecosystem
Google's I/O conference focused on integrating AI across its entire product suite, introducing the Gemini Omni world model and the cost-effective Gemini 3.5 Flash to drive enterprise and consumer adoption.
Introducing Gemini Omni
Google officially introduces Gemini Omni, a new multimodal model designed for real-time, low-latency interaction. This model enhances the capabilities of the Gemini ecosystem by providing seamless cross-modal processing.
Gemma 4 Released: Multimodal Open Models
Google DeepMind launched Gemma 4, open-weights multimodal models in E2B, E4B, 26B A4B, and 31B sizes. They handle text, images, video, and audio, with up to 256K context and strong reasoning, coding, agentic features. Dense and MoE variants optimized for on-device deployment.
Xihu University and Alibaba DAMO Academy Launch Guiyuan AI
Xihu University and Alibaba DAMO Academy have developed 'Guiyuan', an AI model designed to predict stem cell fate. By analyzing millions of molecular combinations, the model enables precise control over cell reprogramming, significantly accelerating biological research.
SenseNova-Vision: Unified Open-Source Visual Foundation Model
SenseTime has released SenseNova-Vision, a unified visual foundation model that integrates tasks like detection, segmentation, and 3D reconstruction. It is fully open-source, including a 50 million sample visual instruction corpus.
Infinity-Parser2: New SOTA Multimodal Document Parsing Model
Infinity-Parser2 is a new multimodal model featuring a 5-million-sample synthetic corpus and a multi-task reinforcement learning framework. It achieves state-of-the-art performance on document parsing benchmarks, outperforming models like DeepSeek-OCR-2.
Meta Releases Muse Spark 1.1 and New Low-Cost API
Meta has unveiled 'Muse Spark 1.1', a multimodal reasoning model with enhanced agentic capabilities for tool use and complex coding. The model is available via the new 'Meta Model API' at a price point lower than competing models.
Meta releases Muse Spark 1.1 for coding
Meta has launched Muse Spark 1.1, an upgraded AI model designed for advanced coding tasks. It features improved bug detection, support for multi-agent workflows, and native multimodal perception.
Tencent Hires OpenAI Researcher Yonglong Tian for VLM
Tencent has hired former OpenAI researcher Yonglong Tian to join its Large Language Model department. He will focus on the research and development of Vision-Language Models (VLM).
Shengshu Tech Launches Vidu S1 Real-time Interactive Model
Shengshu Tech has released Vidu S1, a new video generation model capable of real-time interaction, voice control, and high-definition video output.
Global first: Edge-side streaming multimodal model released
A Hangzhou-based team has successfully deployed a streaming multimodal model to the edge, marking a significant milestone for CVPR 2026 trends.
Pixel-Level Photo Attacks Bypass AI Chatbot Safety Rules
Researchers at Florida International University discovered that invisible pixel-level modifications in images can force AI chatbots to bypass their safety guardrails and generate restricted content.
Introducing Gemma 4 12B: A Unified Multimodal Model
Google DeepMind has released Gemma 4 12B, a new encoder-free multimodal model. This release expands the Gemma family with enhanced capabilities for processing diverse data types.
Google Releases Open-Source Gemma 4 12B Multimodal Model
Google has released the Gemma 4 12B multimodal model, designed to run locally on consumer hardware. It offers performance comparable to the 26B version while requiring only 16GB of memory.
Google Releases Gemma 4 12B Multimodal Model for Local PCs
Google has launched 'Gemma 4 12B', an open multimodal model designed for efficiency. It features an encoder-free architecture that allows it to run on laptops with 16GB of RAM.
World's first omni-modal API now free and open
A top-tier AI lab has released the world's first omni-modal API, which is now available for free indefinitely. The API supports seamless processing of text, images, and video data.
Google Gemini 3.5 and AI Industry Shifts
Google released Gemini 3.5 Flash and Omni models, while the industry sees major shifts with Anthropic hiring Andrej Karpathy and AMD expanding local AI infrastructure. AI competition is evolving from raw parameter counts to ecosystem building and industry penetration.
Google launches Gemini 3.5 and new AI agent products
Google has officially unveiled the Gemini 3.5 model alongside new AI agent capabilities and advanced video generation models. This release marks a significant shift in Google's product strategy toward autonomous AI agents.
Intern-S2-Preview: Efficient 35B Scientific Multimodal Model
Intern-S2-Preview is a new 35B scientific foundation model that scales task difficulty and diversity. It features enhanced agentic capabilities and efficient reinforcement learning using CoT compression.
Ovis2.6-80B-A3B: MoE multimodal model with active vision
AIDC-AI introduced Ovis2.6-80B-A3B, a multimodal LLM using a Mixture-of-Experts architecture. It features 'Think with Image' capabilities, allowing the model to actively manipulate visual inputs during reasoning.
NVIDIA Launches Nemotron 3 Nano Omni Multimodal Model
NVIDIA introduces Nemotron 3 Nano Omni, a compact long-context multimodal model for building intelligent agents handling documents, audio, and video. Hosted on Hugging Face, it advances efficient multimodal AI capabilities. This launch targets agent developers seeking lightweight yet powerful solutions.
SenseTime Open-Sources SenseNova U1 Models
SenseTime officially released and open-sourced the SenseNova U1 series, a native unified model for understanding and generation. Built on the in-house NEO-unify architecture from March, it integrates multimodal understanding, reasoning, and generation. The model treats language and vision as a unified composite for efficient synergy while preserving semantics and pixel fidelity.
Xiaomi MiMo-V2.5 Tops Agentic Claw Efficiency
Xiaomi launched open-source MiMo-V2.5 and V2.5-Pro models under MIT license, available on Hugging Face for commercial use. They lead in agentic 'claw' tasks like OpenClaw, with Pro achieving 63.8% success on ~70K tokens—40-60% fewer than Claude or GPT. Featuring 310B parameters and 1M-token context, they challenge closed-source frontiers.
Qwen3.6-27B Uncensored Aggressive Launches
New fully uncensored Qwen3.6-27B Aggressive model released with optimized K_P GGUF quants. Achieves 0 refusals with no capability loss, supports multimodal inputs. Sensitive to prompt clarity; disable thinking via Jinja kwarg.
Qwen3.6-35B Uncensored Aggressive Released
New uncensored aggressive variant of Qwen3.6-35B-A3B released with K_P quants and full multimodal support. Zero refusals, no capability loss, and optimized for local inference. Includes various quants, vision mmproj, and 262K context.
Claude 4.7 Tops Coding Benchmarks
Anthropic launches Claude Opus 4.7, its top model with 64.3% on SWE-bench Pro (vs GPT-5.4's 57.7%), multi-agent coordination for long workflows, 3x image resolution, and 14% better agentic reasoning with fewer tool errors. Pricing: $5 input/$25 output per million tokens.
Qwen3.6-35B-A3B Open-Source MoE Launched
Qwen3.6-35B-A3B is a new open-source sparse MoE model with 35B total parameters and 3B active params under Apache 2.0. It matches agentic coding of models 10x larger and offers strong multimodal perception with thinking/non-thinking modes. Available on HuggingFace, ModelScope, and Qwen Studio.
Alibaba's Happy Horse Tops Video Gen Benchmarks
Alibaba stealthily released the Happy Horse AI video generator last week. It now ranks top on key video generation benchmarks. This positions Alibaba as the leader in AI video synthesis.
Uncensored Gemma 4 E4B/E2B Multimodal Launch
New aggressive uncensored variants of Gemma 4 E4B (4B) and E2B (2B) released as fully multimodal models supporting text, image, video, and audio. Available in high-quality GGUF quants on Hugging Face, compatible with llama.cpp. Larger E31B and E26B models coming soon.