⚙️
多 GPU KV 快取與專家張量卸載查詢
使用者尋求在 llama.cpp 中指定 KV 快取至特定 GPU 及卸載專家張量的方法。目標是將非專家張量置於強 GPU,其餘置於弱 GPU。已跨貼至 GitHub 討論 #20642。
Tag: #kv-cache63 results
使用者尋求在 llama.cpp 中指定 KV 快取至特定 GPU 及卸載專家張量的方法。目標是將非專家張量置於強 GPU,其餘置於弱 GPU。已跨貼至 GitHub 討論 #20642。

Nvidia's DMS compresses LLM KV cache up to 8x, reducing memory costs without accuracy loss. Enables longer chain-of-thought reasoning and more parallel paths. Outperforms heuristic eviction and paging methods.