Research questionHow can we select language-model populations with low correlated failures when semantic similarity misses generative-process diversity?Models can produce semantically different outputs yet fail on the same tasks because their underlying generative processes may be similar. Semantic similarity alone therefore may not reveal the diversity needed to reduce shared failures.