HeteroFold transfers a sender LLM’s KV cache directly into a different-family receiver LLM without re-running prefill, by aligning tokens on shared character boundaries, remapping key/value features into the receiver’s statistics, and calibrating against the receiver’s native attention. At 32K context it hits about 10.7× speedup over native prefill.
Picture a multi-agent system where a Llama model drafts a long document and then hands it to a Qwen model to critique and a Ministral model to vote. Today each receiver has to re-read the whole document from scratch, running a full Prefill phase pass to rebuild its own KV cache. That is pure duplicated compute, and it scales linearly with context length. At 32K tokens, native prefill into Ministral-3-14B takes about 5 seconds per batch of 2 just to be ready to generate the first token.
The obvious shortcut is to ship the sender’s KV cache over the wire so the receiver can skip prefill. That works when sender and receiver are the same model family (same tokenizer, same layer count, same KV shapes). It breaks the moment the agents are heterogeneous, which is exactly the interesting case for multi-agent setups that mix model sizes and vendors. Prior prefill-free methods like Dense Latent and KV Ridge learn a map from sender KV to receiver KV, but their published evaluations stay inside one family and they assume compatible tokenization. HeteroFold is the first method the authors are aware of that handles the cross-family case end-to-end while keeping both models frozen.
There are three problems to solve in series. First, the two models tokenize the same text differently, so “token 7” in the sender does not correspond to “token 7” in the receiver. Second, they have different numbers of layers and different KV head counts, so you cannot just copy layer-to-layer. Third, even if you reconstruct KV values with low error, the receiver’s attention pattern can still drift. Figure 2 of the paper makes this third point directly: KV Ridge gets lower reconstruction error than HeteroFold but produces worse attention.
Token Alignment (TA). For each receiver token, find the sender token that ends at the same character position in the raw text. If there is no exact match, fall back to the latest earlier shared character boundary. This gives a stable sender index for every receiver position regardless of how the two tokenizers chop up words.
Layer Alignment plus Recolor mapping. For each receiver layer, pick three nearby sender layers (at depth offsets of -4, 0, +4 after proportional depth mapping) and concatenate their K or V features. Then learn an affine map that takes this concatenated sender vector into the receiver’s K or V space. The map is initialized by Recolor: whiten the sender features, rotate them to align with the receiver features using a Procrustes alignment solution, then re-color so the mapped output matches the receiver’s native mean and covariance. This is a closed-form moment-matching initialization, not gradient training.
Output-aware calibration. Even after Recolor, the receiver’s attention still drifts. So they add a small Low-rank correction (rank 16) on top of the affine map and train it against two losses: a KL between native and transferred attention weights (for keys), and a normalized squared error on the attention output after the output projection (for values). Both base models stay frozen; only the rank-16 factors move. After training, the correction is folded back into the affine map, so inference is just one affine transform per layer with no extra module.
# per receiver layer, per role in {K, V}
for rx_tok in receiver_tokens:
sx_tok = latest_shared_char_boundary(rx_tok) # TA
feats = concat([sender_kv[sx_tok, l] for l in nearby_sender_layers])
y = (feats / r) @ A_final + b_final # Recolor + folded correction
kv_cache[rx_tok] = reshape_to_receiver_heads(y)
# receiver then applies its own key-norm and RoPE and decodes
They test six directed transfers among Llama-3.1-8B, Qwen3-4B, and Ministral-3-14B, across four long-context QA tasks (Qasper, HotPotQA, LoCoMo, QuALITY) and five short-context tasks including MMLU, GSM8K, and ARC-Challenge. The reference is TextMas, meaning the receiver just re-prefills the original text, which is lossless but slow.
•
HeteroFold wins on all four long-context benchmarks across all six directions against the prefill-free baselines Dense Latent and KV Ridge (both given the same TA token correspondence for a fair fight). It also wins most short-context settings.
•
Example magnitude: for Ministral-3-14B→Llama-3.1-8B on HotpotQA F1, HeteroFold hits 50.3 versus KV Ridge+TA at 37.5 and TextMas at 57.5. So the method closes most of the gap to lossless text communication without re-prefilling.
•
On HiddenBench, a 15-round multi-agent task where agents must share private facts to find the right answer, HeteroFold reaches 30.6% average accuracy versus TextMas at 29.2% and KV Ridge+TA at 21.8%. This is the headline evidence that the method survives when the transferred content is itself generated during the conversation, not just a fixed prompt.
•
Latency: at 32K context, Llama→Ministral transfer runs in 481 ms versus 5167 ms for native prefill (10.7×) and about 570–705 ms for the other prefill-free baselines (1.18–1.47× over them). HeteroFold is faster partly because it uses one affine map rather than a two-layer MLP, and only three sender layers per receiver layer rather than eight.
•
Ablation (Table 3, Ministral→Llama): removing TA and using same-index token pairing is the single biggest hit, dropping HotpotQA F1 from 50.3 to 0.7. Recolor beats ridge regression as the initialization, and key-side calibration contributes most of the long-context gains.
The authors are careful to note that low KV reconstruction error does not imply correct receiver behavior, which is why they measure attention KL and next-token agreement directly in their analysis section.
•
If you are building a heterogeneous multi-agent pipeline where the same long context (a document, a tool output, a conversation history) gets handed between models, this is a concrete recipe to avoid paying prefill N times. The win scales with context length, so it is most interesting past about 16K tokens.
•
The method needs offline calibration per transfer direction (about two H100-hours per pair under their default: 1,600 training prompts, four epochs). Worth testing if your agent topology is relatively fixed. If you spin up arbitrary new sender/receiver pairs at runtime, that calibration cost matters.
•
You also need the sender’s raw K and V tensors before key-norm and Rotary Position Embedding (RoPE), which means you need the kind of model access you get from running the weights yourself, not from a hosted API. The paper does not establish an equivalent through black-box endpoints.
•
Code is on GitHub. The approach also helps on same-family transfers (their Qwen3 appendix), so it is a reasonable drop-in even when you are not strictly crossing families.
•
As a conceptual takeaway even if you do not deploy it: when evaluating any KV-transfer or KV-compression scheme, measure attention-weight and output divergence directly, not just reconstruction MSE. The paper is a clean demonstration that reconstruction error and downstream behavior can disagree sharply.
•
All results are on three specific models (Llama-3.1-8B, Qwen3-4B, Ministral-3-14B). Scaling to much larger receivers or very different architectures is not shown.
•
Calibration requires a reasonable prompt corpus (they used an equal mix of Open-R1 and HotpotQA) and the sensitivity table shows real degradation at only 200 training prompts.
•
HeteroFold still trails TextMas on several short-context tasks, sometimes by a few points of accuracy. It is a speed/quality trade, not free.
•
HiddenBench is 65 tasks with three or four agents and 15 rounds; the headline parity with TextMas rests on a small eval. Treat the multi-agent result as encouraging rather than conclusive.