NVIDIA Nemotron 3 Ultra 於 Amazon SageMaker JumpStart 正式上線

💡在 AWS 上部署 NVIDIA 最新的推理模型,為代理型 AI 帶來 5 倍推論速度與 30% 的成本節省。
⚡ 30 秒速覽
有什麼變化
現已可在 Amazon SageMaker JumpStart 上進行部署
為什麼重要
此發布降低了企業在生產環境中部署高效能推理模型的門檻。它特別針對代理型 AI 應用所需的效率需求,這對於自動化業務流程而言正變得至關重要。
下一步行動
在 SageMaker JumpStart 上部署一個 Nemotron 3 Ultra 測試實例,以評估您現有的代理型 AI 工作流程在成本與延遲方面的改善。
關鍵要點
- •現已可在 Amazon SageMaker JumpStart 上進行部署
- •針對代理型 AI 工作負載提供 5 倍的推論速度提升
- •相較於先前配置,營運成本降低了 30%
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 15 個來源。
🔑 增強重點摘要
- •NVIDIA Nemotron 3 Ultra is a 550-billion-parameter Mixture-of-Experts (MoE) model with 55 billion active parameters, specifically engineered for frontier reasoning and orchestration in complex agentic AI systems.
- •The model incorporates a sophisticated hybrid Mamba-Attention architecture, leveraging LatentMoE for enhanced accuracy and Multi-Token Prediction (MTP) layers for accelerated inference through native speculative decoding.
- •Nemotron 3 Ultra supports an expansive context length of up to 1 million tokens, which is critical for maintaining coherence and managing extensive information in long-running agentic workflows, such as analyzing large codebases or legal documents.
- •NVIDIA has made the base, post-trained, and quantized checkpoints of Nemotron 3 Ultra, along with its training data and recipes, openly available on HuggingFace, fostering broader customization and development.
- •The model is optimized for integration with leading agent platforms and harnesses, including Hermes Agent, LangChain Deep Agents, OpenClaw, OpenHands, and OpenCode, indicating its design for practical, ecosystem-wide deployment in enterprise AI.
📊 競品分析▸ Show
| Metric/Model | NVIDIA Nemotron 3 Ultra | GLM-5.1-754B-A40B | Kimi-K2.6-1T-A32B | Qwen-3.5-397B-17B |
|---|---|---|---|---|
| Parameters | 550B total / 55B active (MoE) | 754B | 1T | 397B |
| Architecture | Hybrid Mamba-Attention MoE | N/A | N/A | N/A |
| Context Length | Up to 1M tokens | N/A (max 256K for Kimi/Qwen) | N/A (max 256K for Kimi/Qwen) | N/A (max 256K for Kimi/Qwen) |
| Inference Throughput (8K input / 64K output) | 5.9x higher than GLM-5.1, 4.8x higher than Kimi-K2.6, 1.6x higher than Qwen-3.5 | Baseline | Baseline | Baseline |
| Agent Productivity PinchBench | 91% | 84% | 91% | 89% |
| Long-horizon Planning EnterpriseOps-Gym | 33% | 40% | 29% | 30% |
| Coding Terminal-Bench 2.0 | 54% | 64% | 67% | 53% |
| Instruction Following IFBench | 82% | 77% | 74% | 78% |
| Long Context Ruler @1M | 95% | N/A (max 256K) | N/A (max 256K) | 90% |
| Pricing (OpenRouter) | $0 per million input/output tokens | N/A | N/A | N/A |
Note: Detailed feature comparisons and consistent pricing data for competitor models across all platforms were not readily available in the search results. Pricing for Nemotron 3 Ultra on OpenRouter is listed as free, implying infrastructure costs would be the primary expense when deployed on platforms like SageMaker JumpStart.
🛠️ 技術深入
- Parameter Count: 550 billion total parameters with 55 billion active parameters, utilizing a Mixture-of-Experts (MoE) architecture.
- Architecture Type: Hybrid Mamba-Attention Mixture-of-Experts (MoE) with interleaved Mamba-2 and MoE layers, along with select Attention layers.
- Key Technologies: Leverages LatentMoE for improved accuracy and efficient expert routing, and incorporates Multi-Token Prediction (MTP) layers for faster inference through native speculative decoding and improved generative speed.
- Quantization: Pre-trained using NVFP4 quantization, an NVIDIA 4-bit floating point format, to maximize compute efficiency and enable cross-architecture GPU deployment.
- Context Length: Supports an extensive context length of up to 1 million tokens, crucial for long-running and complex agentic tasks.
- Training Pipeline: Post-trained with an enhanced pipeline that includes Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD) for improved model accuracy.
- Optimization: Optimized for agent-led open harnesses and designed to excel in workflows involving planning, tool calling, observation reading, sub-agent delegation, output validation, and error recovery.
- Minimum GPU Requirements: Requires a minimum of 4xGB200, 4xB200, 4x GB300, 4x B300, or 8xH100 GPUs for deployment.
- Supported Languages: Supports English, French, Spanish, Italian, German, Japanese, Korean, Hindi, Brazilian Portuguese, and Chinese.
🔮 前景展望基於引用來源的 AI 分析
⏳ 時間線
📎 來源 (15)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: AWS Machine Learning Blog ↗
每週電子報
每週一封,可隨時退訂。
