Get Started
Home
Topics
Search
Library
Research questionHow can training-data attribution distinguish reweighting influence from genuine intervention leverage?Influence functions estimate effects from small changes in example weights, but such estimates may not predict the behavioral change obtainable by modifying an example. This makes it difficult to separate limited attribution signals from interventions that fail to exploit an example’s leverage.
AI
Alignment & Safety
LLM Pretraining & Post-training
Machine Learning
Latest papersRecent research connected to this question, newest first.From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data AttributionThe evidence concerns influence-function-selected examples in four open-weight LLMs. It compares reweighting with response rewriting while keeping instructions fixed, primarily testing epistemic abstention and also examining safety refusal; conclusions are limited to these settings.research paper · Sep 2, 2026
Related questions
How can we reduce fine-tuning-induced shortcut reliance on underrepresented groups without retraining or group labels?How can we tell whether internal estimates guide effective actions, rather than merely predict effects accurately?How can small-scale pretraining mixture experiments stay reliable when scarce high-quality data is repeated at target scale?How can model distillation block hidden teacher-trait transfer through clean data without degrading the target task?