Get Started
Home
Topics
Search
Library
Research questionHow should multilingual LLMs estimate uncertainty and calibrate abstention across languages and model sizes?Uncertainty estimates can change with the language of the question and reasoning, particularly in low-resource languages, as well as with model scale. These differences make it difficult to identify unreliable answers and set consistent abstention thresholds.
AI
Alignment & Safety
Evaluation & Benchmarks
Natural Language Processing
Reasoning
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMsThe evidence covers multiple-choice question answering in 22 high-, mid-, and low-resource languages, using human-curated datasets and long-form reasoning. It compares nine open- and closed-box uncertainty methods across model sizes and architectures, and analyzes threshold selection for selective prediction; the evaluation avoids LLM-as-a-judge and embedding-based scoring.research paper · Sep 4, 2026
Related questions
How can LLMs distinguish ambiguous inputs from gaps in their knowledge when estimating uncertainty?When does an LLM’s verbal confidence reliably reflect its underlying uncertainty?How can multiple-choice music audio-language models estimate uncertainty well enough to abstain without costly ensembles or retraining?How can LLMs produce reliable confidence estimates for deciding when to defer outputs to humans?