Tencent’s Hy4 Preview Targets Production-Grade AI

💡A 770B model with 1M context, low pricing, and evidence of self-testing code workflows.
⚡ 30-Second TL;DR
What Changed
Hy4 preview grows from Hy3’s 295B total parameters and 21B active parameters to 770B and 49B respectively, with context extended from 256K to 1M tokens.
Why It Matters
Hy4 preview strengthens Tencent’s position in cost-efficient, open-weight-style model access while emphasizing agentic workflows rather than single-turn generation. Its long context, low pricing, and ability to test and revise outputs could make it attractive for enterprise automation and coding agents, although the preview still requires human review of critical assumptions.
What To Do Next
Run a controlled pilot of Hy4 preview through TokenHub or OpenRouter using your longest coding and document workflows, and compare quality, latency, and cost against your current model.
Key Points
- •Hy4 preview grows from Hy3’s 295B total parameters and 21B active parameters to 770B and 49B respectively, with context extended from 256K to 1M tokens.
- •It is available in WorkBuddy, CodeBuddy, Yuanbao, and ima, and can be accessed through Tencent Cloud TokenHub and OpenRouter.
- •Pricing is set at 6 yuan per million input tokens, 18 yuan per million output tokens, and 0.3 yuan per million cached tokens.
- •In Tencent’s blind test of 203 engineering tasks, Hy4 preview scored 2.99/4, slightly ahead of Kimi K3 at 2.94 and GLM 5.3 at 2.92.
- •A hands-on evaluation found it could cross-check expense documents, build a Canvas game with automated tests, and revise implementations after identifying bugs.
🧠 Deep Insight
Background and context from public sources — not the original article. 13 sources cited.
🔑 Enhanced Key Takeaways
- •Tencent released the Hy4 model under the Apache 2.0 license, enabling unrestricted commercial use and community-driven development.
- •The model was trained using a recursive self-improvement loop where Hy4 assisted in optimizing its own training data strategies and evaluation frameworks.
- •Tencent claims a 31.8% increase in end-to-end training throughput as a direct result of the model's contribution to its own development pipeline.
- •Hy4 supports standard open-source inference engines including vLLM and SGLang, facilitating easier enterprise self-hosting compared to proprietary-only models.
- •Tencent provides an FP8 quantized version of the model alongside the full-precision release to lower hardware requirements for production deployment.
📊 Competitor Analysis▸ Show
| Feature | Tencent Hy4 | Moonshot Kimi K3 | Z.ai GLM-5.3 |
|---|---|---|---|
| Architecture | 770B MoE (49B active) | Proprietary | Proprietary |
| Context Window | 1M tokens | N/A | N/A |
| Engineering Score | 2.99/4 | 2.94/4 | 2.92/4 |
| License | Apache 2.0 | Closed | Closed |
| Cache Efficiency | 85% lower cost | Baseline | Baseline |
🛠️ Technical Deep Dive
- Architecture: Mixture-of-Experts (MoE) design with 770B total parameters and 49B active parameters per token.
- Quantization: Native support for FP8 precision to optimize memory footprint and inference speed.
- Compatibility: Full support for vLLM and SGLang inference frameworks for production-grade deployment.
- Training Methodology: Utilized recursive self-improvement loops where the model contributed to its own data strategy and evaluation framework optimization.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (13)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 极客公园 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
