SourceReddit r/LocalLLaMA•Stalecollected in 15h
Atlas Open-Sourced for GB10 Ultra-Fast Inference

Open-source engine hits 130 tok/s on Qwen3.6-35B—3x vLLM on GB10
30-Second TL;DR
What Changed
130 tok/s peak on Qwen3.5-35B NVFP4 MTP K=2
Why It Matters
Democratizes Blackwell inference speeds for community, outperforming vLLM 3x and enabling edge AI on specialized hardware.
What To Do Next
Run 'docker pull avarok/atlas-gb10:latest' and serve Qwen3.6-35B-A3B-FP8 with --speculative flag.
Who should care:Developers & AI Engineers
Key Points
- •130 tok/s peak on Qwen3.5-35B NVFP4 MTP K=2
- •Pure Rust+CUDA, 2.5GB image, no PyTorch overhead
- •Docker deploy supports prefix caching, OpenAI API
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.