Research questionHow can input-adaptive sparse attention reduce long-context prefilling cost without losing retrieval-relevant context?Long-context self-attention prefilling grows quadratically, while fixed sparse patterns can miss input-dependent structure. Dynamic routing can add overhead and misallocate attention mass as context length increases.