Get Started
Home
Topics
Search
Library
Research questionHow many dimensions encode truth in language models, and does that geometry persist for computed rather than retrieved statements?Truth values can be read from language-model hidden states, but it remains unclear whether they are concentrated in one direction or distributed across multiple components. It is also uncertain whether the same representational geometry applies when truth must be computed rather than retrieved.
AI
Machine Learning
Mechanistic Interpretability
Natural Language Processing
Research Paper
Latest papersRecent research connected to this question, newest first.The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Shared Category Geometry in Small Language ModelsThe study examines hidden-state differences for true/false minimal pairs using a training-free dominant SVD direction, evaluated across 14 models from six architectural families, including mixture-of-experts models. It considers whether the observed arrangement extends from retrieved truth to categories whose truth is computed; the extracted direction costs O(d) per token.research paper · Sep 2, 2026
Related questions
How can vision-language models infer 3D geometry and temporal continuity from 2D visual observations?How can multilingual representation sharing be measured without confusing anisotropy with genuine cross-lingual structure?Do pretrained language models encode a reusable truthfulness signal for detecting misinformation without external evidence?How can we distinguish decodable logical validity from reasoning that actually drives a language model’s answers?