Get Started
Home
Topics
Search
Library
6 min read · Evaluation · LLM Training · Added Sep 30 · Paper published Sep 24, 2026

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

Source: research paper via Hugging Face Daily Papers
0:00 / 7:55
Private fine-tunes leak through single-token black-box outputs on prompts unrelated to the task. Querying a teacher at near-tie decisions where the shared public base is ~50/50 between two words, a student trained on 5,664 one-token answers recovers +5.3pp HumanEval+ over a nuisance-matched control.
TL;DR
A private post-training update to a language model leaves a behavioral shadow: it shifts the model’s word choice on prompts that have nothing to do with the target task. A student trained only on those one-word choices recovers part of the teacher’s coding skill.
Why It Matters
Suppose you fine-tune a model on code, then serve it behind a black-box API that only returns the top next token. A competitor sends prompts that have nothing to do with code, records which of two ordinary words your model picks, and uses those picks to make their model better at code. That is the situation this paper constructs and shows is real.
Prior work on Subliminal learning already showed that unrelated teacher outputs can transmit traits like preferring owls, using long teacher generations. This paper pushes that idea in two ways: the observations are single words, and what transfers is a measurable capability (code pass rate on executed tests), not a stylistic preference. The relevant baseline the authors position against is ordinary Knowledge Distillation, which needs teacher probabilities or long teacher responses on task data. Here the student sees neither.
How It Works
The setup assumes a known ancestor: there is a public base model, someone privately fine-tunes it into a teacher, and the student starts from the same public base. The student can query the teacher, but only gets one greedy next token back per query.
The core trick is where to query. For each candidate prompt, the authors use the public base to score two ordinary single-token words (say jacket vs tie). They keep only prompts where the base is within 2% of a coin flip between the two, and where one of those two words is also the base’s overall top token. These are near-tie prompts: places where a tiny nudge to the model’s preferences can flip the output.
Why this matters: think of the private update as a small vector added to the base’s weights. On a near-tie, even a small preference change from that update can reverse which word wins. So the teacher’s one-word answer at a near-tie is a one-bit readout of how the private update moved the model at that particular decision. Across thousands of such prompts, these bits jointly constrain the direction of the update.
The method is called Active Taskless Distillation (Active Taskless Distillation (ATD)). Training is plain cross-entropy on the one retained token per prompt. Nothing in the prompts or the labels mentions the target task, and a frozen audit confirms this for the coding run.
carriers = [] for prompt, a, b in candidate_prompts: q = softmax_pair(base_logits(prompt), a, b) # public base only top1 = argmax(base_logits(prompt)) if abs(q[b] - 0.5) <= 0.02 and top1 in {a, b}: w = greedy_next_token(teacher, prompt) # one black-box call if w in {a, b}: carriers.append((prompt, w)) student = finetune(base, carriers, loss="cross_entropy_on_last_token")
What They Found
The headline experiment uses Qwen2.5-1.5B-Instruct as both the public base and student init, with a private LoRA teacher fine-tuned on code preference data via Direct Preference Optimization. The student trains on 5,664 single-token teacher responses drawn from target-unrelated prompts.
•
On HumanEval+, the student scores 51.22% pass@1, matching the teacher, and beats an exact nuisance-matched control by +5.34 pp. That control reuses the same prompts and the same multiset of teacher tokens but scrambles which token goes with which prompt, holding word frequency and per-difficulty-bin flip counts fixed. So generic fine-tuning on these prompts, or the teacher’s overall word habits, cannot explain the gain. The prompt-to-token correspondence is what carries the signal.
•
Across five independently re-collected acquisitions (fresh prompts, freshly trained teachers) crossed with three training seeds, all 15 signal-minus-control gaps are positive, with a mean of +4.80 pp.
•
The effect is source-specific. A code-trained teacher’s shadow helps most on code; a science-trained teacher’s shadow helps most on science. Off-diagonal effects are near zero or negative. So the channel is not carrying a generic “try harder” signal.
•
The shadows compose. A student trained on a 50/50 mix of near-orthogonal math and code shadows recovers both directions, and the recovered strength tracks the teacher’s own update magnitude across training checkpoints.
•
A big teacher gain is not enough on its own. Two adversarial teachers with huge target-task gaps (one that memorized the HumanEval+ answers, one given a substitution cipher) transfer nothing. The authors read this as: the shadow only moves capability that the student’s ancestor could plausibly express, not arbitrary teacher advantages.
•
The channel is narrow. On MBPP+, where the teacher itself only gains +2.38 pp over the base, the student recovers roughly zero. Where the teacher’s own gain is small, there is little shadow to read.
What’s Useful
For people building or hosting fine-tuned models:
•
Treat single-token black-box outputs as leaking information about your private fine-tune, not just about your base model. Rate-limiting or output-perturbation strategies that assume “one token per query is safe” need rethinking. The paper does not evaluate defenses; this is the implication, not a tested claim.
•
The leak requires the querier to have the same public ancestor your fine-tune started from. If your production model derives from a base the outside world does not have, the demonstrated attack does not directly apply. The authors show two cross-family setups where transfer failed, though they note tokenizer and scale also differed.
For researchers studying distillation or model fingerprinting:
•
Active Taskless Distillation (ATD) is a clean instrument for asking “is capability X observable through decisions on inputs that do not express X.” Worth testing on your own base/fine-tune pairs, especially in regimes where the teacher’s target-task gain is substantial (the paper’s evidence for tiny-gap teachers is negative).
•
The near-tie selection matters, not just the query count. A passive baseline with the same query budget on ordinary prompts underperforms the active version by about 4-5 pp on HumanEval+. If you replicate, do not skip the ancestor-uncertainty filter.
•
Code is available: github.com/myboker/ATD.
Caveats
•
Every positive result requires the student to start from the same public ancestor as the teacher. Cross-family transfer failed in the two tested pairs.
•
The largest confidence intervals on the extension models (Qwen3-1.7B, Qwen3-4B, Llama-3.2-1B) include zero; only the mean is positive. The strong claim is really about the Qwen2.5-1.5B coding lineage.
•
Transfer tracks the teacher’s own gain. On weak-gap benchmarks like MBPP+ the effect vanishes, so “ATD transfers coding capability” should be read as “on benchmarks where the teacher itself improved noticeably.”
•
Two teachers with huge target-task gaps (memorized answers, injected cipher) transferred nothing, so this is not a general answer-laundering attack. It moves capability the student was already latently able to express.
•
All quantitative claims about how much was recovered come from held-out benchmarks (pass@1 on executed tests, exact match, multiple choice). The paper’s separate representational diagnostics measure alignment of internal directions, not additional capability.
Topics
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Evaluation133 episodes
LLM Training134 episodes
NLP92 episodes