Search

Tag: #gguf28 results

Luth-2 Sets a New French SLM Benchmark

Luth-2 Sets a New French SLM Benchmark

Luth-2-0.8B and Luth-2-2B are new non-reasoning French language models designed for local and on-device use. The developers report state-of-the-art results for their size across French benchmarks, with models, GGUF files, training data, and code available now.

Reddit r/LocalLLaMACommunityAug 11#french-language#gguf#post-training
⚙️

Arc Pro B70 Trails RTX 3090 in llama.cpp Benchmarks

Benchmarks compare Nvidia RTX 3090 and Intel Arc Pro B70 using llama.cpp on Vulkan and SYCL backends across multiple GGUF models. Arc Pro B70 shows 71% slower prompt processing and 53% slower token generation on average versus RTX 3090. SYCL backend improves some Arc results over Vulkan.

Reddit r/LocalLLaMACommunityApr 23#benchmarks#vulkan#sycl
Qwen3.6 Autonomously Builds Tower Defense Game

Qwen3.6 Autonomously Builds Tower Defense Game

A user tasked Qwen3.6-35B with building a tower defense game using screenshots from MCP, and it successfully implemented and tested features like upgrades. The model self-detected and fixed bugs in canvas rendering and wave completions. High excitement for the upcoming Qwen Coder model.

Reddit r/LocalLLaMACommunityApr 17#agentic#multimodal#gguf
Gemma 4 31B SpecDec +29% Speedup

Gemma 4 31B SpecDec +29% Speedup

Speculative decoding with Gemma 4 E2B draft model boosts Gemma 4 31B inference by 29% on average, reaching 50% on code generation. Key issue was fixed GGUF metadata mismatch causing token translation overhead. Achieves high acceptance rates on structured tasks like math and code.

Reddit r/LocalLLaMACommunityApr 12#speculative-decoding#benchmark#gguf
⚙️

ZINC: Zig LLM Inference for AMD GPUs

ZINC is a new LLM inference engine written in Zig, enabling 35B models on $550 AMD GPUs via Vulkan. It loads GGUF models, achieves 7.1 tok/s on RDNA4, and addresses AMD consumer GPU gaps. Repo at github.com/zolotukhin/zinc.

Reddit r/LocalLLaMACommunityMar 29#amd-gpu#vulkan#gguf
Page 1 of 3