Research questionHow can we compare feature contributions to language-model representations when those features are correlated?Decoding probes can show that a feature is recoverable from a representation, but they do not directly quantify that feature’s contribution. Correlations among features can further confound interpretation of what a model represents.