Draft-KV: Learning Useful Latent Communication Between Language Models
Draft-KV sends a frozen sharer's draft key-value states to a small receiver, lifting MMLU-Redux from 37.45% to 78.04%.
Draft-KV passes key-value states produced while a frozen sharer drafts an answer, instead of decoded text. Across five earlier method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Linear projections place the states in side memory read by gated attention; both models stay frozen and the interface trains 1.05 million parameters, 348 times fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux versus 37.45% alone and 36.40% with reassigned messages.
- Swapping prior latent messages for unrelated ones changes accuracy by at most 0.60 points.
- Draft-KV sends frozen draft key-value states into gated side memory.
- The interface trains 1.05M parameters, 348 times fewer than C2C.
- A Qwen2.5-0.5B receiver reaches 78.04% MMLU-Redux with a Qwen3-8B sharer, versus 37.45% alone.
Full article185 words · extracted from huggingface.co · click to collapse
Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.34754