來源較早收集於 2m

NVIDIA Nemotron 3 Ultra 於 Amazon SageMaker JumpStart 正式上線

NVIDIA Nemotron 3 Ultra 於 Amazon SageMaker JumpStart 正式上線
PostLinkedIn
☁️閱讀原文: AWS Machine Learning Blog
#aws#agentic-ainvidia-nemotron-3-ultranvidiaamazon sagemakernemotron 3 ultra

💡在 AWS 上部署 NVIDIA 最新的推理模型,為代理型 AI 帶來 5 倍推論速度與 30% 的成本節省。

⚡ 30 秒速覽

有什麼變化

現已可在 Amazon SageMaker JumpStart 上進行部署

為什麼重要

此發布降低了企業在生產環境中部署高效能推理模型的門檻。它特別針對代理型 AI 應用所需的效率需求,這對於自動化業務流程而言正變得至關重要。

下一步行動

在 SageMaker JumpStart 上部署一個 Nemotron 3 Ultra 測試實例,以評估您現有的代理型 AI 工作流程在成本與延遲方面的改善。

誰應關注:Developers & AI Engineers

關鍵要點

  • 現已可在 Amazon SageMaker JumpStart 上進行部署
  • 針對代理型 AI 工作負載提供 5 倍的推論速度提升
  • 相較於先前配置,營運成本降低了 30%

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 15 個來源。

🔑 增強重點摘要

  • NVIDIA Nemotron 3 Ultra is a 550-billion-parameter Mixture-of-Experts (MoE) model with 55 billion active parameters, specifically engineered for frontier reasoning and orchestration in complex agentic AI systems.
  • The model incorporates a sophisticated hybrid Mamba-Attention architecture, leveraging LatentMoE for enhanced accuracy and Multi-Token Prediction (MTP) layers for accelerated inference through native speculative decoding.
  • Nemotron 3 Ultra supports an expansive context length of up to 1 million tokens, which is critical for maintaining coherence and managing extensive information in long-running agentic workflows, such as analyzing large codebases or legal documents.
  • NVIDIA has made the base, post-trained, and quantized checkpoints of Nemotron 3 Ultra, along with its training data and recipes, openly available on HuggingFace, fostering broader customization and development.
  • The model is optimized for integration with leading agent platforms and harnesses, including Hermes Agent, LangChain Deep Agents, OpenClaw, OpenHands, and OpenCode, indicating its design for practical, ecosystem-wide deployment in enterprise AI.
📊 競品分析▸ Show
Metric/ModelNVIDIA Nemotron 3 UltraGLM-5.1-754B-A40BKimi-K2.6-1T-A32BQwen-3.5-397B-17B
Parameters550B total / 55B active (MoE)754B1T397B
ArchitectureHybrid Mamba-Attention MoEN/AN/AN/A
Context LengthUp to 1M tokensN/A (max 256K for Kimi/Qwen)N/A (max 256K for Kimi/Qwen)N/A (max 256K for Kimi/Qwen)
Inference Throughput (8K input / 64K output)5.9x higher than GLM-5.1, 4.8x higher than Kimi-K2.6, 1.6x higher than Qwen-3.5BaselineBaselineBaseline
Agent Productivity PinchBench91%84%91%89%
Long-horizon Planning EnterpriseOps-Gym33%40%29%30%
Coding Terminal-Bench 2.054%64%67%53%
Instruction Following IFBench82%77%74%78%
Long Context Ruler @1M95%N/A (max 256K)N/A (max 256K)90%
Pricing (OpenRouter)$0 per million input/output tokensN/AN/AN/A

Note: Detailed feature comparisons and consistent pricing data for competitor models across all platforms were not readily available in the search results. Pricing for Nemotron 3 Ultra on OpenRouter is listed as free, implying infrastructure costs would be the primary expense when deployed on platforms like SageMaker JumpStart.

🛠️ 技術深入

  • Parameter Count: 550 billion total parameters with 55 billion active parameters, utilizing a Mixture-of-Experts (MoE) architecture.
  • Architecture Type: Hybrid Mamba-Attention Mixture-of-Experts (MoE) with interleaved Mamba-2 and MoE layers, along with select Attention layers.
  • Key Technologies: Leverages LatentMoE for improved accuracy and efficient expert routing, and incorporates Multi-Token Prediction (MTP) layers for faster inference through native speculative decoding and improved generative speed.
  • Quantization: Pre-trained using NVFP4 quantization, an NVIDIA 4-bit floating point format, to maximize compute efficiency and enable cross-architecture GPU deployment.
  • Context Length: Supports an extensive context length of up to 1 million tokens, crucial for long-running and complex agentic tasks.
  • Training Pipeline: Post-trained with an enhanced pipeline that includes Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD) for improved model accuracy.
  • Optimization: Optimized for agent-led open harnesses and designed to excel in workflows involving planning, tool calling, observation reading, sub-agent delegation, output validation, and error recovery.
  • Minimum GPU Requirements: Requires a minimum of 4xGB200, 4xB200, 4x GB300, 4x B300, or 8xH100 GPUs for deployment.
  • Supported Languages: Supports English, French, Spanish, Italian, German, Japanese, Korean, Hindi, Brazilian Portuguese, and Chinese.

🔮 前景展望基於引用來源的 AI 分析

NVIDIA's strategy of open-sourcing advanced models like Nemotron 3 Ultra will accelerate the adoption and innovation of agentic AI across enterprises.
By providing open weights, training data, and recipes, NVIDIA enables broader customization and integration, fostering a more robust ecosystem for AI agent development.
The integration of Nemotron 3 Ultra on Amazon SageMaker JumpStart will significantly lower the barrier to entry for enterprises to deploy sophisticated agentic AI solutions.
The one-click deployment and managed infrastructure offered by SageMaker JumpStart simplify the operational complexities associated with deploying large, complex AI models.
The focus on long-context and efficient agentic workflows with Nemotron 3 Ultra will drive a shift in enterprise AI applications towards more autonomous and multi-step reasoning systems.
The model's capabilities address critical challenges in long-running tasks, such as maintaining coherence and managing costs, which are essential for practical enterprise automation.

時間線

2023-11
NVIDIA introduced Nemotron-3 8B for enterprise chatbot and copilot development.
2025-12
NVIDIA launched the hybrid Nemotron 3 family, including Nemotron 3 Nano.
2026-02
NVIDIA Nemotron 3 Nano 30B model became generally available in Amazon SageMaker JumpStart.
2026-04
NVIDIA Nemotron-3-Super-120B model became available on Amazon SageMaker JumpStart.
2026-04
NVIDIA Nemotron 3 Nano Omni model became available on Amazon SageMaker JumpStart.
2026-06
NVIDIA Nemotron 3 Ultra launched on Amazon SageMaker JumpStart.

📎 來源 (15)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. nvidia.com
  2. amazon.com
  3. nvidia.com
  4. nvidia.com
  5. nvidia.com
  6. ollama.com
  7. startupfortune.com
  8. nvidia.com
  9. openrouter.ai
  10. wikipedia.org
  11. nvidia.com
  12. youtube.com
  13. technewsworld.com
  14. amazon.com
  15. amazon.com
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: AWS Machine Learning Blog

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週電子報

每週一封,可隨時退訂。