Tencent AI accelerates with top-ranked Hy3 preview

💡Tencent's Hy3 is currently dominating OpenRouter usage; see if it fits your production agent stack.
⚡ 30-Second TL;DR
What Changed
Hy3 preview leads OpenRouter in token usage for three consecutive weeks.
Why It Matters
Tencent's aggressive AI agent deployment and model optimization indicate a shift toward high-frequency, specialized AI application delivery in the enterprise market.
What To Do Next
Benchmark your current agent workflows against the Hy3 preview API to evaluate its performance in coding and context-heavy tasks.
Key Points
- •Hy3 preview leads OpenRouter in token usage for three consecutive weeks.
- •Tencent has launched dozens of general and vertical AI agents in 2026.
- •Hunyuan model underwent a full rebuild in under three months to improve coding and context capabilities.
🧠 Deep Insight
Web-grounded analysis with 13 cited sources.
🔑 Enhanced Key Takeaways
- •Tencent's Hy3 preview is a Mixture-of-Experts (MoE) model, featuring 295 billion total parameters with only 21 billion activated per token, and supports an extensive 256K token context window.
- •The recent rebuild of the Hunyuan model's pre-training and reinforcement learning infrastructure, overseen by former OpenAI researcher Yao Shunyu, was specifically geared towards enhancing practical, real-world applications and business needs, moving beyond a sole focus on benchmark scores.
- •Tencent is developing a 'top-secret' AI agent for its WeChat platform, aiming to deeply integrate with its vast mini-program ecosystem to automate complex, multi-step tasks like ride-hailing and food delivery for its 1.4 billion users.
- •The Hy3 preview model demonstrates a 40% improvement in inference efficiency and has proven capable of reliably powering complex agent workflows involving up to 495 steps across various use cases, including document processing and data analysis.
- •Tencent offers the Hy3 preview through its Tencent Cloud TokenHub with competitive pricing, starting at approximately USD 0.18 per million input tokens, indicating a strong push for broad commercial adoption.
📊 Competitor Analysis▸ Show
Competitor Analysis: Tencent Hy3 Preview
| Feature/Metric | Tencent Hy3 Preview | Claude Opus 4.6 | GPT-5.4 | GLM-5 | Kimi-K2 | Llama 3.1 405B (Hunyuan-Large comparison) |
|---|---|---|---|---|---|---|
| Architecture | MoE (295B total, 21B active) | Closed-source, likely MoE | Closed-source, likely MoE | MoE | MoE | MoE (389B total, 52B active) |
| Context Window | 256K tokens | Large | Large | Large | Large | 256K tokens |
| Inference Efficiency | 40% improvement over previous versions | - | - | - | - | Lower processing requirement than dense models of similar size |
| Pricing (Input) | ~$0.18/million tokens (RMB 1.2) | - | - | - | - | Free for developers outside EU with <100M monthly users (Hunyuan-Large) |
| SWE-bench Verified | 74.4% | 80.8% | 78.6% | Competitive | - | - |
| Terminal-Bench 2.0 | 54.4% | - | - | - | - | - |
| WideSearch | 70.2% | - | - | Competitive | Close to Kimi-K2 | - |
| MMLU | - | - | - | - | - | 88.4% (Hunyuan-Large) vs 85.2% (Llama 3.1 405B) |
| AlpacaEval 2 | - | - | - | - | - | 51.8% (Hunyuan-Large-Instruct) vs 50.5% (DeepSeek 2.5 Chat) |
| Agent Capabilities | Powers complex workflows up to 495 steps | - | - | - | - | State-of-the-art agentic performance (Hunyuan-A13B) |
| Key Focus | Production-grade agent workloads, complex reasoning, coding, context learning | Advanced reasoning | Advanced reasoning | General-purpose LLM | General-purpose LLM | General-purpose, open-source, efficiency |
🛠️ Technical Deep Dive
- Model Architecture: Mixture-of-Experts (MoE) model.
- Parameter Count: 295 billion total parameters, with 21 billion parameters activated per token during inference.
- Context Window: Supports a native 256,000 token context window.
- Reasoning Capabilities: Integrates both 'fast thinking' and 'slow thinking' capabilities for enhanced complex reasoning.
- Speculative Decoding: Utilizes a 3.8 billion parameter MTP (Multi-Path Transformer) layer for speculative decoding, aiming for lower latency and smaller batch sizes.
- Internal Structure: Comprises 80 transformer layers, with 192 routed experts (top-8 selection) and 1 shared expert.
- Attention Mechanism: Employs Grouped Query Attention (GQA) with 64 heads over 8 Key-Value (KV) heads.
- Inference Modes: Offers three dynamic inference modes:
no_think,think_low, andthink_high, allowing for a trade-off between latency and reasoning depth. - Infrastructure Rebuild: The model emerged from a rebuilt pre-training and reinforcement learning infrastructure, emphasizing stability and practical business needs.
- Agent Framework Integration: Supports integration with popular open-source agent frameworks such as OpenClaw, OpenCode, and KiloCode.
- Deployment Requirements: To serve on 8 GPUs, requires high-memory GPUs like NVIDIA H200, AMD MI300X/MI325X (192 GB), or AMD MI350X/MI355X (288 GB).
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (13)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪 ↗