Research questionHow can low-bit KV caches save autoregressive decoding memory without losing long-context retrieval quality?Low-bit KV caches reduce memory use but introduce model- and quantizer-dependent quality loss. Improvements in perplexity do not necessarily preserve long-context retrieval. Latest papersRecent research connected to this question, newest first.Quality Recovery for Quantized KV Caches via Low-Rank Attention AdaptationThe evidence evaluates fixed affine, NF4, KIVI, and KVarN cache formats on TinyLlama, Gemma, and Llama models, using low-rank Q/K/V projection adaptation. Results include held-out perplexity, associative retrieval, and a subset of RULER; the reported 2-bit setting recovered perplexity substantially but restored only 11–12 of 180 retrieval cases.research paper · Sep 2, 2026