Get Started
Home
Topics
Search
Library
Research questionHow should reference baselines be chosen to make gradient-based attributions meaningful and reliable?Interpretation results depend on what represents the reference state of an input or feature. An unsuitable or implicit baseline can make attribution values difficult to interpret and can distort evaluations of attribution quality.
AI
Computer Vision
Evaluation & Benchmarks
Machine Learning
Mechanistic Interpretability
Latest papersRecent research connected to this question, newest first.The Neglected Baseline in Model InterpretationThe source analyzes baseline assumptions across gradient-based, integrated-gradient, and Taylor-style interpretations and proposes a revised attribution approach that supports features from different layers. Its evidence is limited to the interpretation methods and evaluation settings examined in the study, along with its attribution-error framing.research paper · Sep 4, 2026
Related questions
How does gradient training build hierarchical representations beyond the lazy kernel regime?How can we test whether reference-based NLG metrics behave correctly under controlled response changes?How can offline preference optimization identify which chosen–rejected pairs merit gradients without destabilizing reasoning-model training?How can language-model attention remain reliable beyond its training context?