Search

Tag: #low-vram11 results

⚙️

vLLM Dynamic Expert Caching for Low-VRAM MoE

A new PR in vLLM introduces dynamic expert caching with LRU policy, enabling 16G MoE models on 8G VRAM by keeping active experts in VRAM and offloading others to RAM. On cache misses, computation shifts to CPU while reshuffling experts to minimize latency. Upcoming features include mxfp4 quantization and disk streaming.

Reddit r/LocalLLaMACommunityMar 17#expert-caching#moe-inference#low-vram
⚙️

Rose: Low-VRAM PyTorch Optimizer Launch

Rose is a new stateless PyTorch optimizer with ultra-low VRAM (less than 8-bit AdamW), fast convergence, and strong generalization, licensed Apache 2.0. Benchmarks on MNIST show competitive accuracy despite sometimes higher training loss but lower validation loss. Easy to use for ML training with minimal memory overhead.

Reddit r/MachineLearningCommunityApr 24#optimizer#low-vram#stateless
Page 1 of 2