Capability-Driven Self-Evolution of Agent Memory
PrisMem evolves agent memory by capability and beats baselines on long-history benchmarks.
Researchers introduce PrisMem, a capability-driven method for evolving executable agent memory programs instead of revising them from mixed overall feedback. It prioritizes capabilities with cross-capability benefits, diagnoses specialists from history, and merges complementary gains using paired differential traces. PrisMem beats the strongest baselines by 10.54 points on BEAM-1M and 7.83 points on LongMemEval-M.
- PrisMem evolves executable agent memory by capability, not overall score.
- Dependency-aware selection and trace-guided integration preserve complementary gains.
- Gains of 10.54 and 7.83 points on BEAM-1M and LongMemEval-M.
- Evaluated on million-token interaction histories.
Full article150 words · extracted from huggingface.co · click to collapse
Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and judging progress by overall performance. This can obscure optimization directions and hide capability-specific gains offset by regressions elsewhere, leaving promising directions underexplored. We introduce capability-driven evolution, which extends search guidance from overall performance to individual capability dimensions, preserving promising revisions and expanding exploration beyond the boundaries of holistic evolution. We propose PrisMem, which uses dependency-aware capability selection to prioritize targets with potential cross-capability benefits and history-guided diagnosis to refine capability specialists. Trace-guided integration compares evaluated programs on paired differential cases, using their behavioral differences to consolidate complementary gains into a unified memory program. Experiments show that PrisMem outperforms the strongest baselines by 10.54 and 7.83 percentage points on BEAM-1M and LongMemEval-M, respectively, demonstrating its effectiveness on million-token histories.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.06361