Get Started
Home
Topics
Search
Library
Research questionHow can we select language-model populations with low correlated failures when semantic similarity misses generative-process diversity?Models can produce semantically different outputs yet fail on the same tasks because their underlying generative processes may be similar. Semantic similarity alone therefore may not reveal the diversity needed to reduce shared failures.
AI
Alignment & Safety
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Inferred Generative-Process Diversity Predicts Correlated Failure Across Language ModelsThe evidence covers 38 language models and uses raw model outputs to infer generative-process diversity with Normalized Compression Distance residualized against a permutation control. Across ten disjoint benchmark families, the measure is associated with chance-corrected correlated failure beyond semantic similarity and model-pair capability; the findings establish an association rather than a causal guarantee for population selection.research paper · Sep 3, 2026
Related questions
How can we measure whether generative models preserve conditional diversity without multiple reference outputs per input?How can we detect internally incoherent language-model forecasts before relying on them for consequential decisions?How can we build reproducible Armenian LLMs without sacrificing knowledge for fluency?How can models adapt to low-resource languages without damaging source-language and related-task performance?