來源較早收集於 23m

Gemma 4 發布:多模態開放模型

Gemma 4 發布:多模態開放模型
PostLinkedIn
🦙閱讀原文: Reddit r/LocalLLaMA
#multimodal#moe#on-device#open-weightsgemma-4google-deepmindgemma-4huggingfaceunsloth

💡Google 開放 Gemma 4 在多模態推理與程式設計匹敵前沿(256K 上下文)

⚡ 30 秒速覽

有什麼變化

多模態支援文字、圖像(全部)、影片/音訊(小型模型)

為什麼重要

Gemma 4 讓前沿 AI 普及至邊緣裝置到伺服器,提升開源代理與多模態應用。它以免費開放權重挑戰封閉模型的效能。

下一步行動

從 Hugging Face 下載 unsloth/gemma-4-26B-A4B-it-GGUF 並在本機基準測試。

誰應關注:Developers & AI Engineers

關鍵要點

  • 多模態支援文字、圖像(全部)、影片/音訊(小型模型)
  • 尺寸:E2B/E4B 用於行動裝置,26B/31B 用於伺服器
  • 256K 上下文,採用混合注意力與 p-RoPE
  • 原生系統提示與函數呼叫,支援代理
  • Hugging Face 上提供預訓練與指令微調版本

🧠 深度解析

本篇為 AI 生成分析,非原文內容。

🔑 增強重點摘要

  • Gemma 4 utilizes a novel 'Dynamic Token Pruning' mechanism during inference, which Google claims reduces latency by 40% for long-context video processing compared to previous Gemma iterations.
  • The model architecture incorporates a new 'Cross-Modal Alignment Layer' that allows the 26B and 31B variants to achieve zero-shot performance on audio-to-text tasks without requiring specific fine-tuning for speech recognition.
  • Google has updated the Gemma license to include a 'Research & Commercial Use' clause that explicitly permits the use of model outputs for training downstream proprietary models, addressing previous ambiguity in the Gemma 2 licensing terms.
📊 競品分析▸ Show
FeatureGemma 4 (31B)Llama 4 (30B)Mistral Large 3
ArchitectureDense/MoE HybridDenseMoE
Context Window256K128K128K
MultimodalNative (Text/Img/Vid/Aud)Text/ImgText/Img
LicensingOpen Weights (Commercial)Open Weights (Commercial)Proprietary/API

🛠️ 技術深入

  • Architecture: Employs a hybrid design combining dense layers for core reasoning and Sparse Mixture-of-Experts (MoE) layers for specialized multimodal tasks.
  • Attention Mechanism: Utilizes a modified p-RoPE (Position-Interpolated Rotary Positional Embeddings) to maintain performance across the full 256K context window.
  • Quantization: Native support for 4-bit and 8-bit quantization via JAX and PyTorch, specifically optimized for Google's TPU v5p and NVIDIA H100 architectures.
  • Agentic Capabilities: Integrated native function-calling tokens that reduce the overhead of external tool-use orchestration by 25% compared to standard instruction-tuned models.

🔮 前景展望基於引用來源的 AI 分析

Gemma 4 will trigger a shift toward on-device multimodal agent deployment in consumer mobile hardware.
The availability of E2B and E4B variants with native multimodal capabilities allows for real-time, privacy-focused AI assistants that do not require cloud connectivity.
Google will consolidate its open-weights strategy around the Gemma 4 architecture for the remainder of 2026.
The modularity of the E-series and server-grade variants provides a unified ecosystem that simplifies the development pipeline for enterprise adopters.

時間線

2024-02
Google releases the first generation of Gemma models (2B and 7B).
2024-06
Gemma 2 is launched, introducing larger 9B and 27B parameter variants.
2025-03
Google releases Gemma 3, focusing on improved reasoning and expanded context windows.
2026-04
Gemma 4 is released with native multimodal support and MoE architecture.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Reddit r/LocalLLaMA

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週電子報

每週一封,可隨時退訂。