Get Started
Home
Topics
Search
Library
Research questionHow can bias evaluations of large language models diagnose affected groups and reasons behind biased outputs?LLM bias assessments can produce constrained results that are difficult to interpret. Without diagnostic detail, it is hard to identify which demographic groups are affected or why an output is judged biased.
AI
Alignment & Safety
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.GPTBIAS: A Comprehensive Framework for Evaluating Bias in Large Language ModelsThe source describes GPTBIAS, which uses a high-performing LLM such as GPT-4 to assess bias in other models with Bias Attack Instructions. It reports bias scores alongside bias types, affected demographics, keywords, reasons, and improvement suggestions, with experiments addressing effectiveness and usability.research paper · Sep 2, 2026
Related questions
How can subtle media bias be detected without task-specific annotated training data?How susceptible are large language models to conspiratorial responses under demographic conditioning?How can safety evaluations measure language-model behavior without triggering evaluation-aware changes in decisions?How can evaluators distinguish missing knowledge from miscalibrated outputs in language models?