πŸ¦™Stalecollected in 25m

Krasis Hits 8.9x Prefill Speed vs Llama.cpp

Krasis Hits 8.9x Prefill Speed vs Llama.cpp
PostLinkedIn
πŸ¦™Read original on Reddit r/LocalLLaMA

πŸ’‘4-9x faster than llama.cpp on single GPU for 122B+ models

⚑ 30-Second TL;DR

What Changed

8.9x prefill, 4.7x decode vs llama.cpp on 5090

Why It Matters

Enables consumer GPUs like 5080/5090 to run massive models faster than llama.cpp, lowering barriers for high-performance local inference.

What To Do Next

Install Krasis via single GitHub command and benchmark Qwen3.5-122B on your 5090.

Who should care:Developers & AI Engineers

Key Points

  • β€’8.9x prefill, 4.7x decode vs llama.cpp on 5090
  • β€’GPU-only for prefill/decode, minimal CPU/RAM needs
  • β€’Runs Qwen3.5-122B-A10B at 2897 tok/s prefill, 27.7 tok/s decode
  • β€’Supports Qwen3.5-235B-A22B; single-line GitHub install
πŸ“°

Weekly AI Recap

Read this week's curated digest of top AI events β†’

πŸ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.