Get Started
Home
Topics
Search
Library
Topic · 6 recaps

Mechanistic Interpretability

Reverse-engineering neural networks at the level of circuits and features — figuring out what specific weights and activations compute, not just what the model outputs.
Sort
Newest
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
Evaluation · Aug 28
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Mechanistic Interpretability · Aug 13
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
Evaluation · Aug 2
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
Image Generation · Jul 21
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
Mechanistic Interpretability · May 28
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Evaluation · May 7
— End of list —