Research questionHow can we distill capable small language models with fewer training tokens without losing teacher behavior?Distillation can lose information when teacher weights are compressed, representations are poorly aligned, or informative feed-forward activations are ignored. The resulting small model must retain the teacher’s behavior while requiring substantially less training compute and data.