Research questionHow can LLMs produce reliable confidence estimates for deciding when to defer outputs to humans?LLM outputs may sound certain even when they are unreliable, making it difficult to set a dependable threshold for human intervention. Confidence estimates must correspond meaningfully to actual correctness, including when labels have ordinal relationships.