Inverse Distillation Unlearning (IDU) trains a fast one-step image generator from a slow Diffusion model or Flow matching teacher while suppressing a chosen “forget set,” by mixing forget samples with the student’s own outputs and asking the generator to supply only whatever is missing to reconstruct the teacher’s distribution.
Modern image generators like diffusion and flow models produce great samples but need many sampling steps, and they can memorize training images you later want gone: a copyrighted artwork, a specific face, an NSFW photo. Two separate problems usually get solved separately. First, people distill the slow multi-step teacher into a one-step student generator for cheap inference. Second, people run machine unlearning to scrub specific samples from the trained model, typically by fine-tuning with a “forget loss” that pushes the model away from those samples and a “remain loss” that keeps quality on everything else.
That two-loss setup is adversarial: one loss degrades the forget data, the other defends the rest, and they fight. It also assumes you still have the retained training data on hand. And almost all of this literature targets the multi-step teacher, not the one-step student people actually ship. The one prior attempt at unlearning for one-step generators the authors compare against, UOT-Unlearn (UOT-Unlearn (Choi et al., 2026)), relies on a separate feature extractor and an Optimal transport assignment cost rather than a teacher. The gap this paper fills: do both jobs at once (distill + unlearn), for a one-step student, without needing the retained data or an auxiliary classifier.
Start with the intuition. There is a known trick called inverse distillation: given a frozen teacher that was trained to match some data distribution, you can run its training objective backwards to recover which distribution the teacher was fit to. Concretely, you train a student distribution and an auxiliary “fake” model in a min-max game; the game bottoms out exactly when the student distribution equals the teacher’s original training distribution. That is how you normally distill a one-step generator from a diffusion teacher.
IDU’s move: instead of letting the student try to recover the whole teacher distribution, pretend the recoverable distribution is already a known mixture: a fixed fraction \rho of the forget samples (which you have) plus (1-\rho) of whatever the student generates. If the teacher was originally trained on (forget + retained) data and the mixture is forced to contain the forget part explicitly, then the only way for the min-max game to close is for the student to fill the slot with the retained distribution. The forget component is “already accounted for,” so the generator stops wasting capacity reproducing it. The paper proves this: at \rho equal to the true forget fraction, the optimal student exactly equals the retained distribution.
A single scalar \rho controls everything. There is no separate forget loss adversarially pulling against a remain loss; both the forget data and the generated data feed into one coherent objective via the “fake” model, which is a scratchpad network that learns to match the current mixture so the generator can be scored against the teacher. For the generator update the authors reuse the Score identity Distillation (SiD) trick with a scaling constant in [0.5, 1.2].
for step in range(N):
x_gen = G(sample_latent(B)) # student samples
x_fgt = sample(forget_set, B) # known forget samples
t, noise = sample_time_and_noise(B)
xg_t, xf_t = add_noise(x_gen, t, noise), add_noise(x_fgt, t, noise)
# 1) fake model learns the rho-mixture of forget + generated
loss_fake = rho * mse(f(xf_t,t), target(x_fgt)) \
+ (1-rho) * mse(f(xg_t,t), target(x_gen))
update(f, loss_fake)
# 2) generator pushed so (fake on mixture) matches frozen teacher on gen samples
loss_gen = sid_loss(teacher(xg_t,t), f(xg_t,t), target(x_gen), alpha_sid)
update(G, loss_gen)
The generator and the fake model alternate. Training needs only the frozen teacher and the forget samples. No retained images, no external classifier, no feature extractor.
Experiments use MNIST and CIFAR-10 with two teacher families: a flow-matching U-Net and the EDM-VP score-based model used by SiD. The main test forgets two classes per dataset (digits {3,7}; automobile+truck on CIFAR-10). Quality on kept classes is measured with Fréchet Inception Distance (FID) against retained-only data, and forgetting with FGR, the percentage of 50,000 generated images that an off-the-shelf classifier assigns to a forgotten class. Note: class labels define the forget set and evaluate FGR, but IDU itself only ever sees forget samples, not labels.
•
Forgetting is strong. Across the four backbone/dataset combos, per-class FGR stays at most 0.56% for flow matching and at most 1.25% for SiD, versus roughly 10% for the un-unlearned teachers (near the natural class frequency).
•
Retention is competitive with an oracle. Retain-FID for IDU is close to a “retrain-from-scratch-on-retained-data then distill” baseline in three of four settings, and better than that oracle on MNIST flow matching (though the authors flag the MNIST oracle as unstable).
•
Broad across classes. Repeating single-class forgetting across all ten classes gives mean FGR around 0.16%\u20130.84% depending on setup, so the main-table numbers are not cherry-picked pairs.
•
\rho is a real knob. On CIFAR-10 flow matching, \rho=0.6 is the sweet spot; \rho=0.9 diverges. For SiD, \rho=0.2 (which matches the “2 of 10 classes” theoretical value) works best; pushing \rho to 0.9 restores forgotten-class generation to near teacher levels while hurting FID, so bigger is not better.
•
Versus the closest prior one-step unlearner. Against UOT-Unlearn’s reported numbers on comparable CIFAR-10 single-class setups, IDU’s Retain-FID and FGR are both lower (better) on classes 1, 6, 8.
Two honest wrinkles. For flow matching, IDU suppresses the target classes early but FGR can creep back up after ~20k generator updates, so you must checkpoint-select. This reversal was not seen for SiD within the evaluated horizon.
•
If you ship a one-step image generator distilled from a diffusion or flow teacher, and you get a takedown request for specific images (not a whole class concept), this is the first method aimed squarely at that use case. The prerequisite is real: you need the original multi-step teacher checkpoint, not just the deployed one-step student. If you only have the student, this method does not apply as described, though the paper shows fine-tuning the pure-distilled student with the same objective also works on CIFAR-10 at some trade-off cost.
•
Start \rho at the forget-set fraction and treat it as the only tuning knob worth sweeping. The ablations show both too low (weak forgetting) and too high (quality collapse without extra forgetting) fail predictably. The authors used \rho \in {0.2, 0.4, 0.6} for their main runs.
•
If you want to replicate on your stack, SiD-style backbones looked more stable in these experiments than the flow-matching one, which needed careful checkpoint selection. Worth testing before committing.
•
For class-level or concept-level erasure (prompted by a label or text, not sample-defined), the paper notes a different line of work (Score Forgetting Distillation (SFD) is the closest comparator) is still the natural starting point, though IDU reaches comparable numbers on class-defined CIFAR-10 setups without ever seeing class labels during training.
•
Code and checkpoints are promised on publication; the paper does not provide a current link in the supplied text.
•
Only MNIST and CIFAR-10, only unconditional generation, only two teacher families. No text-to-image, no faces, no large-scale demonstration. Scaling behavior is unknown.
•
Needs the trained teacher. Pure black-box API access to a one-step generator is not enough.
•
Needs the forget samples themselves. Concept-level or prompt-defined erasure is out of scope.
•
Flow-matching runs showed late-training reversal of forgetting; the fix is checkpoint selection, which assumes you can evaluate FGR, which assumes you have a classifier for the forget concept. For arbitrary “forget this face” use cases, building that evaluator is itself nontrivial and the paper does not address it.
•
Evaluating Retain-FID requires retained data; the paper suggests substituting teacher samples filtered by classification when retained data is unavailable, but does not quantify how well that substitute tracks the true metric.