Get Started
Home
Topics
Search
Library
Research questionHow can we determine whether recurring multilingual SAE features causally control translation across languages?A sparse autoencoder feature may recur across multilingual discovery settings without representing the same translation mechanism. The difficulty is distinguishing recurrence from a feature whose amplification or ablation reliably changes translation behavior.
AI
Evaluation & Benchmarks
Mechanistic Interpretability
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3Applies to SAE features in Gemma 2 and Gemma 3, tested through activation amplification and ablation across 23 language settings. The evidence reports many recurring features with small or inconsistent effects and one consistently causal feature in each model, without establishing generality beyond these models and settings.research paper · Sep 4, 2026
Related questions
How can we test whether language models causally use scientific mechanisms instead of answer-correlated shortcuts?How can multilingual representation sharing be measured without confusing anisotropy with genuine cross-lingual structure?How can multilingual question answering remain consistent across languages without erasing culturally appropriate differences?How can sparse autoencoder features be shared across language models without per-model retraining?