Get Started
Topic · 9 recaps
Mechanistic Interpretability
Reverse-engineering neural networks at the level of circuits and features — figuring out what specific weights and activations compute, not just what the model outputs.
Play all
...
Posts
Questions
Home
Topics
Search
Library
Sort
Newest
Steering Geometry: Validating Human Value Geometry in LLM Steering Space
Alignment · Sep 5
0
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
Mechanistic Interpretability · Sep 4
0
A Decodable Feature Is an Audit Lead, Not a Capability Contract
Evaluation · Sep 3 · 14:55
0
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
Evaluation · Aug 28
0
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Mechanistic Interpretability · Aug 13 · 7:33
0
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
Evaluation · Aug 2 · 6:59
0
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
Image Generation · Jul 21 · 6:22
0
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
Mechanistic Interpretability · May 28 · 6:59
0
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Evaluation · May 7 · 7:10
0
— End of list —