ByteDance Volcengine Announces Limited-Time AI Plan Discounts

💡Access top-tier models like DeepSeek V4 and MiniMax M3 at a fraction of the cost.
⚡ 30-Second TL;DR
What Changed
Discounts apply to Coding Plan and Agent Plan subscriptions until August 27, 2026.
Why It Matters
Aggressive pricing from Volcengine increases competition in the Chinese enterprise AI market, making high-end model access more affordable for developers and startups.
What To Do Next
Evaluate the Volcengine Coding Plan if you are looking for a cost-effective, integrated environment for AI-assisted development.
Key Points
- •Discounts apply to Coding Plan and Agent Plan subscriptions until August 27, 2026.
- •Integrated models include MiniMax M3, DeepSeek V4, and GLM-5.1.
- •Pricing starts at 9.9 RMB/month for the first two months, down from 40 RMB.
- •Agent Plan includes proprietary Doubao-Seed models for multi-modal tasks.
🧠 Deep Insight
Web-grounded analysis with 28 cited sources.
🔑 Enhanced Key Takeaways
- •Volcengine's discount strategy is part of a broader push to capture a larger share of China's public cloud LLM market, where it already held a 46.4% market share by token calls as of mid-2025, surpassing Baidu AI Cloud and Alibaba Cloud combined.
- •The integrated models offer advanced capabilities: MiniMax M3 features a 1-million-token context window and native multimodal understanding (image and video), excelling in coding and agentic tasks with a proprietary MiniMax Sparse Attention (MSA) architecture.
- •DeepSeek V4, another featured model, is a 1-trillion-parameter Mixture-of-Experts (MoE) model with a 1-million-token context window, native multimodal support (text, images, video, audio), and is positioned as a cost-effective competitor to Western frontier models.
- •GLM-5.1, also included, is a 754-billion parameter MoE model with a 202K context window, designed for long-horizon agentic engineering, demonstrating strong performance in coding and autonomous task execution over extended periods.
- •The Agent Plan specifically leverages proprietary Doubao-Seed models, such as Doubao-Seed-2.0-lite, which offers full-modal understanding (video, images, audio, text), enhanced agentic capabilities with improved multi-turn instruction compliance, and integrated GUI understanding and execution.
📊 Competitor Analysis▸ Show
markdown
| Feature/Provider | ByteDance Volcengine (Agent/Coding Plan) | Baidu AI Cloud (DuClaw) | Tencent Cloud (Agent Development Platform) |
|---|---|---|---|
| Pricing (Base) | 40 RMB/month (Agent Plan Basic), 9.9 RMB/month (discounted for 2 months) | 17.8 RMB/month (introductory rate) | Varies by tier (Free, Starter, Team, Enterprise); Hunyuan model prices increased over 450% in March 2026 |
| Key Models | MiniMax M3, DeepSeek V4, GLM-5.1, Doubao-Seed models | OpenClaw agent platform | Hunyuan series, GLM 5, MiniMax 2.5, Kimi 2.5 (previously) |
| Core Offering | Integrated Agent Plan with multimodal models, web search, Vision Embedding, AFP billing | Zero-deployment AI agent service for OpenClaw | Agent development and management platform with NLP, multi-language, dialog management, and RAG+LLM capabilities |
| Context Window | Up to 1M tokens (MiniMax M3, DeepSeek V4), 202K tokens (GLM-5.1), 256K tokens (Doubao-Seed) | Not explicitly detailed for DuClaw, but OpenClaw compatible models support large contexts | Varies by model, e.g., Doubao-seed-1.6 supports 256K contexts |
| Multimodality | Native multimodal (MiniMax M3, DeepSeek V4, Doubao-Seed-2.0-lite) | Not explicitly detailed for DuClaw, but OpenClaw can integrate vision | Supports multi-modal understanding (e.g., Doubao-seed-1.6) |
| Benchmarks | MiniMax M3: 59.0% on SWE-Bench Pro; DeepSeek V4 Pro: ~91.2% on SWE-Bench Verified; GLM-5.1: 58.4% on SWE-Bench Pro | Not specified for DuClaw directly, relies on underlying models | Not explicitly detailed for ADP, but Hunyuan models have benchmarks |
🛠️ Technical Deep Dive
- MiniMax M3: Employs a proprietary MiniMax Sparse Attention (MSA) architecture, which reduces per-token compute to one-twentieth of the prior generation at 1-million-token context length, enabling over 9x faster prefill and 15x faster decoding. It is natively multimodal, trained from scratch on interleaved text, image, and video data.
- DeepSeek V4: Features a Mixture-of-Experts (MoE) architecture with 1.6 trillion total parameters and approximately 49 billion active parameters per token. It utilizes a hybrid attention architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to achieve significant efficiency gains (27% of single-token inference FLOPs and 10% of KV cache compared to V3.2 at 1M-token context). It supports native multimodal input (text, images, video, audio) and offers three reasoning modes: Non-think, Think High, and Think Max.
- GLM-5.1: A 754-billion parameter Mixture-of-Experts (MoE) model with 40 billion active parameters per token and a 202,752-token context window. Its core innovation is an iterative reasoning framework that allows the model to sustain optimization over hundreds of rounds, enabling autonomous work on complex tasks for up to 8 hours by continuously revisiting reasoning, running experiments, and revising strategies.
- Doubao-Seed Models: The Doubao-Seed-2.0-lite is a full-modal understanding model capable of native unified understanding of video, images, audio, and text. It features upgraded Agent and Coding capabilities, improved compliance with multi-turn complex instructions, stronger self-decomposition and verification, and integrated GUI understanding and execution, allowing it to perform operations like clicking and dragging on interfaces.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (28)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- tmtpost.com
- medium.com
- 36kr.com
- datanorth.ai
- medium.com
- ollama.com
- openrouter.ai
- deepseek.ai
- nvidia.com
- mindstudio.ai
- deepinfra.com
- bentoml.com
- analyticsvidhya.com
- siliconflow.com
- z.ai
- deepinfra.com
- howaiworks.ai
- aibase.com
- zenmux.ai
- wedoany.com
- investing.com
- biggo.com
- tencentcloud.com
- gartner.com
- wandb.ai
- substack.com
- binance.com
- ai-supremacy.com
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: IT之家 ↗