Kimi K3 Hits 92 Tokens/s on Eight B300 GPUs

๐กSee why faster B300 inference beats cheaper A100 hosting on cost per token.
โก 30-Second TL;DR
What Changed
Eight B300 GPUs on Modal delivered 92 tokens/s steady decode and 0.92โ1.02 seconds TTFT.
Why It Matters
The benchmark shows that lower hourly GPU pricing does not necessarily produce lower cost per token when throughput collapses. For large-model serving, high-end GPU throughput and quantization support may matter more than raw instance price.
What To Do Next
Benchmark your target workload with vLLM on the available GPU tier and compare cost per output token rather than hourly instance price.
Key Points
- โขEight B300 GPUs on Modal delivered 92 tokens/s steady decode and 0.92โ1.02 seconds TTFT.
- โขThe deployment used vLLM, tensor parallelism across eight GPUs, native MXFP4, and required about 27 minutes to cold boot.
- โขThe 1-bit UD-IQ1_S GGUF fit on eight A100-80GB GPUs at 9 tokens/s, with 7โ60 seconds TTFT.
- โขB300 inference cost about $190 per million output tokens versus roughly $620 for the slower A100 GGUF setup.
๐ง Deep Insight
Background and context from public sources โ not the original article. 12 sources cited.
๐ Enhanced Key Takeaways
- โขKimi K3 is a 2.8-trillion-parameter Mixture-of-Experts (MoE) model that activates 16 out of 896 experts per token.
- โขThe model architecture incorporates Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) to optimize long-context processing.
- โขMoonshot AI officially recommends a minimum of 64 accelerators for production-grade serving, significantly higher than the 8-GPU setup used in the Reddit experiment.
- โขNota AI has introduced 'Non-uniform Expert Pruning' techniques, which can reduce the hardware footprint of Kimi K3 from eight B300 GPUs down to four or six.
- โขThe model features a mandatory, non-disableable 'thinking' process, with all reasoning tokens billed at the standard output rate of $15 per million tokens.
๐ Competitor Analysisโธ Show
| Feature | Kimi K3 (Self-Hosted) | Claude Opus 4.8 | GPT-5.6 Sol |
|---|---|---|---|
| Weights | Open-Weights | Proprietary | Proprietary |
| Context Window | 1M Tokens | 2M Tokens | 1.5M Tokens |
| Pricing (Output) | ~$190/1M (B300) | $15.00/1M (API) | $15.00/1M (API) |
| Architecture | 2.8T MoE | Proprietary | Proprietary |
๐ ๏ธ Technical Deep Dive
- Model Scale: 2.8 trillion parameters total.
- MoE Configuration: 896 total experts, 16 active experts per token.
- Memory Footprint: Approximately 1.56 TB of VRAM required for model weights.
- Quantization: Native MXFP4 support via vLLM.
- Context Window: Native 1-million-token support.
- Optimization: Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) for long-horizon reasoning.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (12)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
