Research questionHow can we semantically compare autonomous-driving image subsets at scale and attribute their differences to objects?Metadata and fixed labels can indicate what images contain but often cannot explain how two subsets differ semantically, while manual inspection does not scale. Sparse differences also need to be linked to particular object instances or categories to make dataset composition interpretable.