Research questionHow can reasoning benchmarks distinguish models that fail on human-hard items from those failing on human-easy ones?Aggregate benchmark scores can hide whether a model struggles with problems that people also find difficult or makes surprising errors on problems people usually solve. A human-grounded difficulty signal is therefore needed to interpret model failures more precisely.