Research questionHow can 70B language models fit on one GPU while preserving long-context speed and accuracy?A 70B model must fit its weights and growing KV cache within one GPU’s limited memory. Long prompts make compression choices affect both decoding speed and model accuracy.