Research questionHow can we select visual instruction-tuning examples under a fixed budget while preserving alignment and dataset coverage?Visual instruction-tuning corpora are growing quickly, but training on only a budgeted subset can discard examples that support alignment and instruction following. Selection must account for both the coherence of individual visual instruction-response samples and coverage of the broader data distribution.