Search

Tag: #memory-efficiency7 results

Apple's Stochastic KV Routing Cuts Cache Memory

Apple's Stochastic KV Routing Cuts Cache Memory

Apple proposes Stochastic KV Routing to enable adaptive depth-wise cache sharing in transformer models, reducing KV cache memory for high-throughput serving. It optimizes the depth dimension, orthogonal to temporal compression methods. Prior research indicates full per-layer caches are redundant.

Apple Machine LearningOfficialMay 5#kv-cache#memory-efficiency
VibeVoice 1.5B Runs Locally on iPhone

VibeVoice 1.5B Runs Locally on iPhone

A community implementation demonstrates VibeVoice 1.5B running locally on an iPhone with about 2.2 GB of memory and speeds up to 1.28× real time. The model has been uploaded to the audio.cpp Hugging Face repository, with an xcframework and code branch planned after audio.cpp 0.6.

Reddit r/LocalLLaMACommunityAug 5#mobile-ai#on-device#long-form-audio