Search

Few direct matches — filled in with the latest updates.

Tag: #rust-llm1 results

Atlas Unlocks 102 tok/s on DGX Spark

Atlas Unlocks 102 tok/s on DGX Spark

Atlas, a pure Rust LLM inference engine with custom CUDA kernels for GB10's SM121 architecture, delivers 102 stable tok/s on Qwen3.5-35B-A3B, 2.3x faster than vLLM. It features 2-minute cold starts, tiny 2GB image, and excels on Qwen3-Next-80B-A3B at 82 tok/s. Solves DGX Spark's software issues for desktop petaflop compute.

Reddit r/LocalLLaMACommunityMar 4#inference-engine#rust-llm#moe-optimization