Research questionHow can we robustly compress LLM KV caches across open-domain inputs without input-specific budget thresholds?KV-cache pruning can reduce inference memory, but the budget that preserves performance may depend on each input or domain. Open-domain workloads vary substantially, making a fixed threshold unreliable.