Krasis Hits 3324 tok/s Prefill on RTX 5080

๐กNew runtime runs 80B MoE at 3k+ tok/s prefill on one 5080
โก 30-Second TL;DR
What Changed
GPU prefill at 3,324 tok/s on RTX 5080 for 80B MoE
Why It Matters
Enables practical local inference of huge MoE models on consumer GPUs, slashing prefill times for IDE/tools.
What To Do Next
Download Krasis from source and benchmark Qwen3-Coder-Next Q4 on your NVIDIA GPU.
Key Points
- โขGPU prefill at 3,324 tok/s on RTX 5080 for 80B MoE
- โขCPU decode with 14.9 tok/s on Qwen3-Coder-Next
- โขRequires 2.5x model size in system RAM
- โขTargets MoE models, NVIDIA-only, BF16 safetensors input
๐ง Deep Insight
Background and context from public sources โ not the original article. 7 sources cited.
๐ Enhanced Key Takeaways
- โขRTX 5080 uses Blackwell architecture with 10,752 CUDA cores and 16GB GDDR7 VRAM, enabling high AI inference speeds but limited to smaller or quantized models due to VRAM constraints.[1][2]
- โขIn general LLM benchmarks, RTX 5080 achieves around 135 tok/s with two loaded models and up to 26.1 tok/s in specific Vulkan tests, far below Krasis's 3324 tok/s prefill.[4]
- โขRTX 5080 outperforms RTX 6000 Ada in Mistral (4635 vs 4255) and Llama2 (4790 vs 3957) tests but trails RTX 5090 and RTX 4090 in most AI workloads.[1]
๐ฎ Future ImplicationsAI analysis grounded in cited sources
๐ Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- storagereview.com โ Nvidia Geforce Rtx 5080 Review the Sweet Spot for AI Workloads
- microcenter.com โ Benchmarking AI on Nvidia 5080
- youtube.com โ Watch
- youtube.com โ Watch
- pugetsystems.com โ Nvidia Geforce Rtx 5090 Amp 5080 AI Review
- ordinarytech.ca โ How Much Vram Will Games Really Use in 2026 Rtx 5070 Ti vs 5080 vs 5090 Explained 1
- youtube.com โ Watch
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.