KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding
KeyRec lets VLMs answer long-video questions using 10% of dense visual tokens via bounded memory.
KeyRec is a training-free method that bounds visual memory for streaming and long-video vision-language models. It stores fine-grained recent observations in a visual cache and compresses older evidence into an event bank updated online by add, merge, and evict. A text-only router then splits a fixed readout budget between recent and event memory. Across four benchmarks and three backbones, it is best in 13 of 15 compressed settings using 10% of the dense decoder-facing visual-token budget, and it leads every NEO-ov 2B setting.
- Training-free KeyRec keeps recent frames in a visual cache and history in an event bank.
- A text-only router splits a fixed readout budget without reprocessing old frames.
- Best compressed result in 13 of 15 settings at 10% of the dense token budget.
- Gains of 2.21–18.37 points on real-time questions versus the strongest compressed baseline.
- Works on modular encoder–projector VLMs and encoder-free NEO-ov 2B.
Full article207 words · extracted from huggingface.co · click to collapse
Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but may lose coherent event evidence and fail to distinguish detailed recent observations from long-range history. We propose KeyRec, a training-free framework for constructing bounded visual memory. During query-agnostic writing, KeyRec preserves fine-grained recent observations in a visual cache and organizes historical evidence into a structured event bank. Candidate events are proposed according to their novelty relative to previously stored events and maintained through an online add--merge--evict update. When a question arrives, a text-only router adaptively allocates a fixed readout budget between recent and event memory, without reprocessing historical frames. KeyRec operates on model-facing visual embeddings and supports both modular encoder--projector VLMs and the encoder- and projector-free NEO-ov architecture. Across four streaming and long-video benchmarks and three VLM backbones, KeyRec achieves the best compressed performance in 13 of 15 settings using only 10\% of the dense decoder-facing visual-token budget. It outperforms the strongest compressed baseline by 2.21--18.37 points on real-time questions, achieves the best compressed result in five of six long-video settings, and performs best in every NEO-ov 2B setting.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.32182