Get Started
Research questionHow can model distillation block hidden teacher-trait transfer through clean data without degrading the target task?A teacher’s hidden bias can create subtle preference gaps in otherwise ordinary training outputs. Supervised fine-tuning may accumulate updates from those gaps, causing the student to adopt the teacher’s behavior even when the data reveals no obvious trait.
Alignment & Safety
LLM Pretraining & Post-training
Machine Learning
Latest papersRecent research connected to this question, newest first.Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT DistillationThe paper links transfer to trait-direction drift and evaluates a targeted regularization defense during distillation. Evidence includes malicious-response transfer and animal-preference transfer experiments in the main Qwen setting, with reported reductions in transfer and low main-task accuracy cost.research paper · Sep 2, 2026
Related questions
How can contrastive distillation efficiently transfer representations to smaller students without memory banks or fixed temperatures?How can one deployable LLM learn from multiple specialized teachers when the most reliable teacher varies by sample?How can autoregressive logit distillation preserve local next-token preferences when teacher and student rankings disagree?How can on-policy distillation prevent early student errors from corrupting rollouts in autoregressive vision-language models?
Home
Topics
Search
Library