Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
HeteroFold transfers KV caches across model families without prefill, speeding Llama-to-Ministral 32K inference 10.7×.
HeteroFold is a prefill-free method for transferring key-value caches between heterogeneous LLM families while keeping sender and receiver frozen. It aligns model structure, maps the sender cache into the receiver’s representation, and calibrates it so receiver behavior is preserved. Across six transfer directions it leads four long-context benchmarks and most short-context settings, and matches text-based communication on a multi-agent benchmark. At 32K context, Llama-3.1-8B to Ministral-3-14B transfer is 10.7× faster than native prefill and 1.18–1.47× faster than Dense Latent and KV Ridge.
- HeteroFold transfers KV caches across frozen, different model families.
- It aligns structure, maps cache space, and calibrates receiver behavior.
- Best results on four long-context benchmarks across six transfer directions.
- Llama-3.1-8B to Ministral-3-14B at 32K is 10.7× faster than native prefill.
Full article159 words · extracted from huggingface.co · click to collapse
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose HeteroFold, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8BrightarrowMinistral-3-14B transfer is 10.7times faster than Native Prefill and 1.18--1.47times faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.32259