
Atlas Unlocks 102 tok/s on DGX Spark
Atlas, a pure Rust LLM inference engine with custom CUDA kernels for GB10's SM121 architecture, delivers 102 stable tok/s on Qwen3.5-35B-A3B, 2.3x faster than vLLM. It features 2-minute cold starts, tiny 2GB image, and excels on Qwen3-Next-80B-A3B at 82 tok/s. Solves DGX Spark's software issues for desktop petaflop compute.





