Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
Researchers store encrypted KV cache on NVMe so Gemma 4 recalls facts across 50 million tokens without recompute.
The paper evaluates galahad-kv, which writes each roughly 16,000-token block's key-value state to encrypted local NVMe and reloads it byte-exact instead of recomputing. Tests used vLLM on one NVIDIA H100 with Gemma 4 12B and Gemma 4 31B over 50 million tokens of public text. All 100 probed blocks reloaded without recompute; loads were 2.8x to 4.3x faster and used 8.8x to 12.3x less GPU energy, while GPU memory stayed flat. Planted-fact recall was 82 of 100 for the 12B model and 98 of 100 for the 31B model, with no fabricated answers. Only one stored block is loaded at a time, and the store requires terabytes of disk.
- galahad-kv saves encrypted KV blocks of about 16,000 tokens to local NVMe.
- Tested on Gemma 4 12B and 31B via vLLM on a single H100.
- Block reload was 2.8–4.3x faster and used far less GPU energy than recompute.
- Fact recall at multi-million-token depth was 82% (12B) and 98% (31B).
- This reuses stored state one block at a time, not a wider attention window.
Full article265 words · extracted from huggingface.co · click to collapse
A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B. Every block we probed was loaded back from the encrypted store with no recompute (100 of 100, at depths from 0 to 50M tokens) on both models. Loading a block was 2.8x to 4.3x faster than recomputing it and used 8.8x to 12.3x less GPU energy, and GPU memory stayed flat over the whole 50M-token stream. Asked about facts planted millions of tokens earlier, the 12B model gave the right answer 82 times out of 100 and the 31B model 98 times out of 100. Neither model made up an answer. The limits are as follows. This is reuse of stored state, not a wider attention window: one block is loaded at a time, and how well a question is answered depends on the model. Writing the memory is a one-time cost, and the store takes terabytes of local NVMe disk. We describe the test protocol, which is built to resist common ways of gaming long-context benchmarks, and give a single-GPU reproduction that uses public software and a free licence for the package.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.10845