🦙Stalecollected in 65m

6-GPU K80 Multiplexer: 0.3ms Model Hot-Swaps

6-GPU K80 Multiplexer: 0.3ms Model Hot-Swaps
PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#gpu-multiplexing#cheap-vram#kernel-hack#model-switching6-gpu-k80-multiplexernvidiak80rwkv-xbtc-s37

💡$200 for 72GB VRAM + 0.3ms model swaps on old K80s—perfect for local LLM hackers.

⚡ 30-Second TL;DR

What Changed

72GB VRAM from 3x K80 cards (~$200 total)

Why It Matters

Enables ultra-cheap high-VRAM inference rigs for local LLM experimentation, reviving legacy hardware. Democratizes fast multi-model switching for resource-limited builders.

What To Do Next

Source BTC-S37 motherboards on eBay and test K80 multiplexing for cheap multi-model inference.

Who should care:Developers & AI Engineers

Key Points

  • 72GB VRAM from 3x K80 cards (~$200 total)
  • 0.3ms average switch time between 6 dies
  • Custom Linux kernel module for single PCIe multiplexing
  • 38 tok/s decode on RWKV-X 0.2B (INT8)
  • Pure C inference engine, no Python dependencies

🧠 Deep Insight

Background and context from public sources — not the original article. 9 sources cited.

🔑 Enhanced Key Takeaways

  • NVIDIA Tesla K80 features two GK210 GPUs with 4992 CUDA cores total and compute capability 3.7, limiting support to CUDA 11.8 and earlier[2][5][6].
  • Each K80 GPU provides 12GB GDDR5 memory at 480 GB/s aggregate bandwidth, enabling the 72GB total across six dies[1][2][3].
  • K80 offers strong double-precision performance up to 2.91 TFLOPS per card with GPU Boost, advantageous for certain scientific workloads[1][2][4].
  • Recent repurposing efforts highlight K80 viability for offline ML on older CUDA versions, avoiding Tensor Core dependencies[6][8].

🛠️ Technical Deep Dive

  • K80 dual-GPU board uses two GK210 dies, each with 2496 CUDA cores, 12GB GDDR5, and Kepler architecture (compute capability 3.7)[2][5].
  • Supports CUDA up to version 11.8; CUDA 12+ drops Kepler support, requiring compatible LLM frameworks without Tensor Cores[6].
  • Features GPU Boost for dynamic clock scaling, double shared memory/register file vs predecessors, and 480 GB/s memory bandwidth[1][3][4].
  • Both GPUs on K80 can be utilized simultaneously in frameworks like MATLAB for parallel computing[7].

🔮 Future ImplicationsAI analysis grounded in cited sources

K80 multiplexing enables low-cost VRAM scaling until 2028 CUDA support ends
Kepler's CUDA 11.8 limit allows continued use with legacy frameworks, but post-2028 drops will force migration to newer hardware.
Sub-ms switching democratizes multi-model serving on surplus datacenter GPUs
Custom kernel multiplexing leverages cheap K80s for efficient inference, potentially inspiring similar hacks for other EOL cards.

Timeline

2014-11
NVIDIA launches Tesla K80 as dual-GPU Kepler accelerator with 24GB GDDR5.
2022-12
CUDA 12 released, dropping Kepler (K80) architecture support.
2026-03
Reddit r/LocalLLaMA posts 6-GPU K80 multiplexer achieving 0.3ms model hot-swaps on BTC-S37.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.