When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
Pre-registered tests show raw turns selected by Jev match extracted agent memory at far lower cost.
A pre-registered study compared agent memory on held-out LoCoMo conversations and LongMemEval. At a tight budget, raw turns selected by one call to Jev, a typed decision model, were non-inferior to an LLM-extraction memory (one-sided 95% bound −3.0 points against a −5-point margin) and cost 3,061 times less to write. Reranking added 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates were kept, shrinking to 1.5 and 1.1 at generous budgets, where extraction was more accurate. At matched context, Jev matched an LLM reranker (bound −2.0) at one-third the latency and beat multi-call graph traversal.
- On a tight LoCoMo budget, Jev-selected raw turns were non-inferior to LLM-extracted facts.
- Writing raw turns cost 3,061 times less than extraction; the result held with a second answer model.
- Reranking added 17.4 and 9.1 points at three of 30 candidates, but only 1.5 and 1.1 at generous budgets.
- At matched context, Jev matched an LLM reranker at one-third the latency and beat graph traversal.
- Reranking lowered correct abstention; plans, code, and graded answers were released.
Full article201 words · extracted from huggingface.co · click to collapse
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.34227