Get Started
Home
Topics
Search
Library
Research questionHow can we evaluate LLM agents’ moral coherence without shared standards—preserving verdicts under irrelevant changes and responding to morally relevant ones?Moral verdicts may change when wording changes even though morally relevant features remain fixed, while coherent behavior should respond when those features change. This complicates alignment evaluation without relying on a normative reference or expert baseline.
AI
AI Agents
Alignment & Safety
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent AlignmentThe evidence covers three simulated deployments involving nine frontier LLMs, with five paraphrases, five escalation levels, and three dominance conditions. It evaluates verdict stability, monotonicity, decisiveness, and Pareto viability from behavior alone, reporting no model with coherent behavior across all deployments and verdict shifts of up to 99 percentage points at one escalation level.research paper · Sep 4, 2026
Related questions
How can we assess whether LLM moral decisions are defensible when no ground truth exists?How can LLMs give moral advice without being swayed by one-sided multi-turn narratives?How does moral-dilemma context shape human and LLM judgments of autonomous agents’ moral agency?How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?