Search

Tag: #inference-engine11 results

⚙️

ZINC: Zig LLM Inference for AMD GPUs

ZINC is a new LLM inference engine written in Zig, enabling 35B models on $550 AMD GPUs via Vulkan. It loads GGUF models, achieves 7.1 tok/s on RDNA4, and addresses AMD consumer GPU gaps. Repo at github.com/zolotukhin/zinc.

Reddit r/LocalLLaMACommunityMar 29#amd-gpu#vulkan#gguf
Atlas Unlocks 102 tok/s on DGX Spark

Atlas Unlocks 102 tok/s on DGX Spark

Atlas, a pure Rust LLM inference engine with custom CUDA kernels for GB10's SM121 architecture, delivers 102 stable tok/s on Qwen3.5-35B-A3B, 2.3x faster than vLLM. It features 2-minute cold starts, tiny 2GB image, and excels on Qwen3-Next-80B-A3B at 82 tok/s. Solves DGX Spark's software issues for desktop petaflop compute.

Reddit r/LocalLLaMACommunityMar 4#inference-engine#rust-llm#moe-optimization
Page 1 of 2