Get Started
Home
Topics
Search
Library
Research questionHow can we compare feature contributions to language-model representations when those features are correlated?Decoding probes can show that a feature is recoverable from a representation, but they do not directly quantify that feature’s contribution. Correlations among features can further confound interpretation of what a model represents.
AI
Audio & Speech
Audio & Speech Processing
Machine Learning
Mechanistic Interpretability
Natural Language Processing
Research Paper
Latest papersRecent research connected to this question, newest first.Beyond Decodability: Reconstructing Language Model Representations with an Encoding ProbeThe evidence comes from evaluations of text and speech transformer models with feature sets spanning acoustics, phonetics, syntax, lexicon, and speaker identity. Reported effects include strong variation in speaker-related representations across training objectives and datasets, alongside independent syntactic and lexical contributions to reconstruction.research paper · Sep 3, 2026
Related questions
How can sparse autoencoder features be shared across language models without per-model retraining?How can we measure whether transformer representations distinguish senses of the same word across contexts?How can multilingual representation sharing be measured without confusing anisotropy with genuine cross-lingual structure?How can we compare autoregressive and masked-diffusion language models without conflating formulation with architecture?