Get Started
Research questionWhen can plausible but unfaithful LLM self-explanations still support sound decisions?LLM-generated rationales can sound convincing while failing to reveal the processes that produced an answer. This makes it difficult to know whether they should inform decisions, even when they appear useful.
AI
Evaluation & Benchmarks
Natural Language Processing
Reasoning
Latest papersRecent research connected to this question, newest first.From Plausible to Actionable: A Position on LLM Self-ExplanationsThe source is a position paper on natural-language self-explanations generated by LLMs. It critiques standard XAI evaluation practices and discusses actionability across stakeholders, but does not provide empirical evidence establishing when such explanations are reliable.research paper · Sep 7, 2026
Related questions
How can we assess whether LLM moral decisions are defensible when no ground truth exists?How can LLMs give moral advice without being swayed by one-sided multi-turn narratives?How can users judge whether an individual LLM recommendation merits reliance without objective ground truth?How can we evaluate LLM agents’ moral coherence without shared standards—preserving verdicts under irrelevant changes and responding to morally relevant ones?
Home
Topics
Search
Library