
Qwen3.6 27B GGUF Usable for Complex Coding
Qwen3.6-27B-UD-Q6_K_XL.gguf runs at 50 tok/s on RTX 5090 with 200k context via llama.cpp. Excels on difficult planning tasks, promising vs prior local models.
Few direct matches — filled in with the latest updates.
Tag: #llamacpp4 results

Qwen3.6-27B-UD-Q6_K_XL.gguf runs at 50 tok/s on RTX 5090 with 200k context via llama.cpp. Excels on difficult planning tasks, promising vs prior local models.
Delta-KV quantizes KV cache deltas instead of absolutes, achieving near-lossless 4-bit compression on Llama 70B with 10,000x less error. Integrated into llama.cpp fork, adds 10% decode speed via weight-skip. Hardware-agnostic, no training needed.
Add -np 1 to llama.cpp command for solo use to slash SWA KV cache VRAM by 3x (e.g., 3200MB to 1200MB on 31B). Recent PR fix ensures SWA quantizes with KV. Avoid high -ub like 4096 to prevent bloat.
Qwen 3.5 models like 35B A3B require bf16 KV cache in llama.cpp for accurate perplexity, not default f16. Tests show bf16 PPL at 6.5497 vs 6.5511 for f16/f32. vLLM defaults to bf16 correctly, but llama.cpp needs manual flags.

OpenAI and Meta are seeking help to address growing public opposition to their AI data center plans. The effort highlights the increasing public-relations challenges surrounding large-scale AI infrastructure expansion.

VB Pulse data shows enterprises are adopting multiple AI orchestration platforms instead of relying on a single vendor. The survey also highlights persistent concerns about security, permissions, token usage, cost visibility, and the ability to stop runaway agent spending in real time.

US senators are pressing TikTok for answers about an experiment that reportedly withheld an algorithm safety feature. A teenager who later died by suicide was reportedly included in the test.

At 23, Sean Grindal is building around the belief that AI tools increasingly fail because they are difficult to use, not because they lack capabilities. His approach focuses on making AI video creation more accessible as competitors continue adding features.

A new Pew Research Center study finds signs of AI authorship across roughly one-third of web pages published since ChatGPT launched in late 2022. The finding highlights the rapid spread of AI-assisted content creation across the open web.

Google is introducing a button that lets readers mark publishers as preferred sources across Search, Discover, and Google News. The feature could help publishers recover traffic as AI-powered search generates fewer direct clicks to websites.