Research questionHow can on-device LLM inference overlap weight I/O and computation under tight DRAM without stale sparsity decisions?Activation-sparsity decisions become more accurate with later context, while effective I/O-computation overlap requires decisions early. This tension can serialize execution or cause redundant weight transfers and excessive cache use.