Research questionHow can weight-inheritance distillation train the full deployed student matrix rather than a restricted slice?A student may deploy a full-width MLP matrix while optimization can vary only a teacher-induced slice of that matrix. The unreachable parameters consume inference capacity but cannot improve the student, potentially limiting distillation efficiency.