Research questionWhen do compute-saving responsible-AI evaluations preserve full-benchmark conclusions?Changing batching, numerical precision, or benchmark size alters the measurement protocol, so stable aggregate accuracy may not imply stable bias, subgroup, or reasoning conclusions. Reduced subsets can also be sensitive to which benchmark items are retained.