LOCI: Spatial Linear Memory for Streaming World Models
LOCI uses geometry-conditioned spatial memory so video world models recall revisited scenes with less peak memory.
LOCI is a hybrid spatial-memory architecture for streaming video world models. Half of the transformer blocks keep a key-value cache of past observations, while the other half uses recurrent linear attention whose reads and writes are conditioned on projective camera geometry. On the public MIND memory benchmark and held-out recorded trajectories, it reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model, lowering peak memory about 30% with full history. With a bounded bank of retained observations it streams long videos at constant memory and stays more faithful than full softmax under the same budget.
- Hybrid of a key-value cache and geometry-conditioned linear memory
- Camera geometry conditions both memory addressing and stored content
- About 30% lower peak memory than full softmax at equal length
- Streams long video at constant memory with a bounded observation bank
Full article209 words · extracted from huggingface.co · click to collapse
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.40222