Tencent Open-Sources 770B Hunyuan Hy4

💡A 1M-context open model with low token pricing could change the economics of coding and document AI.
⚡ 30-Second TL;DR
What Changed
Model size increased from 295B to 770B total parameters, with active parameters rising from 21B to 49B.
Why It Matters
Hy4 preview gives developers a relatively low-cost option for long-context production workloads, especially coding and enterprise document analysis. Its open-source availability and access through OpenRouter could increase competition among Chinese and global model providers.
What To Do Next
Prototype a long-context coding or document-analysis workflow through OpenRouter, then compare Hy4 preview’s quality, latency, and token cost with your current model.
Key Points
- •Model size increased from 295B to 770B total parameters, with active parameters rising from 21B to 49B.
- •The context window expanded from 256K to 1M tokens for larger codebases and document workflows.
- •Pricing is 6 yuan per million input tokens, 18 yuan per million output tokens, and 0.3 yuan per million cached tokens.
- •Tencent claims inference throughput improved 31.8% versus its baseline after optimizing training methods, data strategies, evaluation, and underlying operators.
🧠 Deep Insight
Background and context from public sources — not the original article. 11 sources cited.
🔑 Enhanced Key Takeaways
- •The model is released under the Apache 2.0 license, facilitating broad commercial and research adoption across platforms like Hugging Face and ModelScope.
- •Tencent utilized a 'model-in-the-loop' development strategy where the model actively participated in refining its own training data strategies and inference system architecture.
- •Internal blind evaluations involving 163 experts across 203 engineering tasks showed Hy4 outperforming competitors GLM-5.3 and Kimi K3.
- •The architecture incorporates a native 10B-parameter Multi-Token Prediction (MTP) layer specifically engineered to accelerate speculative decoding performance.
- •The model is designed for specialized vertical integration, including the capability to generate functional, playable game prototypes directly from natural language prompts.
📊 Competitor Analysis▸ Show
| Feature | Hunyuan Hy4 | GLM-5.3 | Kimi K3 |
|---|---|---|---|
| Architecture | 770B MoE (49B active) | Proprietary | Proprietary |
| Context Window | 1M Tokens | N/A | N/A |
| Expert Eval Score | 2.99/4.00 | 2.92/4.00 | 2.94/4.00 |
| License | Apache 2.0 | Proprietary | Proprietary |
🛠️ Technical Deep Dive
- Architecture: Mixture-of-Experts (MoE) with 770B total parameters and 49B active parameters.
- Attention Mechanism: Utilizes Gated DeepSeek Sparse Attention (DSA) for efficient long-context processing.
- Residual Streams: Implements identity Hyper-Connections (iHC) to stabilize training at scale.
- Speculative Decoding: Features a native 10B-parameter Multi-Token Prediction (MTP) layer to boost inference speed.
- Throughput Optimization: Achieved 31.8% improvement via co-design of training methods and underlying operator kernels.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (11)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 极客公园 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.