SourceStalecollected in 15h

Atlas Open-Sourced for GB10 Ultra-Fast Inference

Read original on Reddit r/LocalLLaMA
#open-source#blackwell#speculative-decode

Open-source engine hits 130 tok/s on Qwen3.6-35B—3x vLLM on GB10

30-Second TL;DR

What Changed

130 tok/s peak on Qwen3.5-35B NVFP4 MTP K=2

Why It Matters

Democratizes Blackwell inference speeds for community, outperforming vLLM 3x and enabling edge AI on specialized hardware.

What To Do Next

Run 'docker pull avarok/atlas-gb10:latest' and serve Qwen3.6-35B-A3B-FP8 with --speculative flag.

Who should care:Developers & AI Engineers

Key Points

  • 130 tok/s peak on Qwen3.5-35B NVFP4 MTP K=2
  • Pure Rust+CUDA, 2.5GB image, no PyTorch overhead
  • Docker deploy supports prefix caching, OpenAI API

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.