Research questionHow can bias evaluations of large language models diagnose affected groups and reasons behind biased outputs?LLM bias assessments can produce constrained results that are difficult to interpret. Without diagnostic detail, it is hard to identify which demographic groups are affected or why an output is judged biased.