Get Started
Research questionHow can we robustly compress LLM KV caches across open-domain inputs without input-specific budget thresholds?KV-cache pruning can reduce inference memory, but the budget that preserves performance may depend on each input or domain. Open-domain workloads vary substantially, making a fixed threshold unreliable.
AI
AI Memory
Inference Optimization
Natural Language Processing
Latest papersRecent research connected to this question, newest first.ReFreeKV: Towards Threshold-Free KV Cache CompressionThis applies to KV-cache pruning for LLM inference. The source reports experiments across 13 datasets spanning context lengths, task types, and model sizes; evidence is limited to those evaluated settings.research paper · Jun 26, 2026
Related questions
How can long-context LLM inference reduce KV-cache memory without losing attention-head-specific information?How can LLM serving adapt KV-cache capacity as attention demand changes during long-output reasoning?How can long-reasoning KV caches retain useful context while reducing memory and eviction overhead?How can large language models cut training and inference costs without materially harming accuracy?
Home
Topics
Search
Library