Can KV Cache Become a Searchable Memory?
The discussion proposes treating a transformer’s KV cache as a high-dimensional, navigable vector space rather than a flat array. This could enable indexing and localized attention, reducing the need to scan all stored context at every inference step.



