Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds
EvolvingNav tracks moving targets so robots can navigate worlds that change while unobserved.
EvolvingNav addresses navigation when targets move while unobserved, including during the agent's trip. It builds a time-indexed belief from timestamped 3D object histories, distinguishing persistence, relocation, and locations outside the known set, then updates that belief with visibility-conditioned RGB-D evidence. A frozen zero-shot vision-language controller uses the belief to choose actions and replan. On EvoWorld-Bench, with 54 scenes and 803,680 tasks, and in real-robot tests, it improves success and search efficiency over baselines, especially when temporal patterns are learnable.
- EvolvingNav predicts target locations that change while unobserved.
- Its belief separates persistence, relocation, and unknown places.
- Visibility-aware RGB-D evidence revises hypotheses without reuse.
- EvoWorld-Bench contains 54 scenes and 803,680 tasks.
- Gains are largest when temporal movement patterns are learnable.
Full article238 words · extracted from huggingface.co · click to collapse
Persistent spatial memory enables embodied agents to navigate familiar environments across repeated visits. However, targets may move while unobserved, including during navigation, making remembered locations unreliable by the time an agent arrives. Despite advances in memory retrieval and state prediction, accounting for continued hidden world evolution and revising beliefs under limited visibility remain challenging. We study Evolving-World Navigation, where agents infer target locations from intermittent observations, predict their states at inspection time, and revise beliefs using visual evidence. We propose EvolvingNav, which constructs a time-indexed belief from timestamped 3D object histories through a structured persistence-relocation model. The belief distinguishes persistence at the last observed location from relocation to alternative locations and retains probability mass outside the known candidate set. An event-driven filter propagates the current belief as time elapses, forecasts target occupancy at candidate inspection times, and incorporates new RGB-D evidence. Negative observations downweight location hypotheses according to calibrated, visibility-conditioned detection probabilities, while evidence tracking prevents repeated use of the same observations. A frozen, zero-shot vision-language controller uses the updated belief to choose actions and replan. We further introduce EvoWorld-Bench, a benchmark grounded in human activity traces, comprising 54 scenes and 803,680 tasks with controlled changes before and during navigation. In simulation and real-robot experiments, EvolvingNav improves navigation success and search efficiency over the evaluated baselines. Paired experiments show the clearest gains under learnable temporal patterns, while ablations demonstrate the value of preserving uncertainty and incorporating visibility-aware evidence.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.39166