Search

Tag: #llama-cpp14 results

⚙️

50 t/s Qwen 3.6 27B on RTX 3090

Tutorial shares setup for 50 tokens/second on RTX 3090 using MTP GGUF Qwen3.6-27B with llama.cpp PR #22673 and 100k context. Includes exact server command with Q4 cache, spec-type mtp, and optimizations. MAC support via mtplx noted.

Reddit r/LocalLLaMACommunityMay 6#mtp-gguf#llama-cpp
Page 1 of 2