Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost
Galahad caches byte-exact KV state so repeated LLM document reading becomes a one-time cost.
Galahad is a memory layer for vLLM, SGLang, and llama.cpp that stores exact key-value state so repeated document text is not recomputed. On seven datasets, 98.7% of prompt tokens were text the model had already read. With Gemma 4 31B and a 97,000-token corpus, Taliesin answered 98 of 100 recall questions versus 10 without it; adding Blaise reached 100 of 100 at about 0.6 seconds and 200 joules, beating a tuned RAGFlow pipeline at 77.
- 98.7% of prompt tokens on seven datasets were already-read text.
- Taliesin restores bit-identical KV state and fails closed on bad loads.
- Blaise cut context to about 668 tokens and beat RAGFlow 100 to 77.
- Galahad worked with all 30 models tested under vLLM.
Full article321 words · extracted from huggingface.co · click to collapse
A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work the model has already done. Serving is stateless across requests: a model that answers a second question about a document recomputes the document's attention state from the first token. On seven real-world datasets, 98.7% of prompt tokens were text the model had already read. We present Galahad, a memory layer for vLLM, SGLang and llama.cpp that makes this reading a one-time cost. Taliesin saves the model's key-value (KV) state for a block of text and loads it on the next request that contains the same bytes, instead of recomputing it. Blaise keeps the documents themselves and passes the model only the section a question needs. On a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone let the model attend to the whole corpus and answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question, against 10 of 100, 9.3 s and 2,754 J for the same model without Galahad, which could hold only the last 12,000 tokens. With Blaise added, the model read about 668 tokens per question and answered 100 of 100 on all three runtimes at 0.59-0.64 s and 200-213 J; a tuned RAGFlow pipeline answered 77. Storing the corpus is a one-time cost of about 100 s and 28 kJ, whose energy is recovered after 13 questions. Restored state is bit-identical: all 262,144 output logits matched after restart, rehydration and hot-load. Galahad worked with all 30 models we tested under vLLM, and it fails closed: any load that does not pass its checks is recomputed. Together these results move LLM serving from stateless to stateful inference.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.39358