Foresight: planning future perception in streaming VLMs without retraining
FORESIGHT lets frozen Qwen3-VL-8B plan future perception in streaming video without retraining.
FORESIGHT is a training-free dual-stream architecture that lets streaming vision-language models anticipate near-future context and reconfigure computation online. Two weight-shared Siamese LLMs share encoders and KV cache: one processes the live stream while the other plans reasoning timing, checks, and sampling density. With frozen Qwen3-VL-8B, it reaches 23.0 mean joint F1 on OmniPro Online, 9.5 points above the strongest trained baseline, and improves StreamingBench by 6.7 and OVO-Bench by 15.4, with a largest gain of 18.7 when evidence arrives late.
- Dual-stream Siamese LLMs share weights, encoders, and KV cache
- Plans when to reason, what to check, and how densely to sample
- Frozen Qwen3-VL-8B scores 23.0 joint F1, beating the trained baseline by 9.5
- Gains 6.7 on StreamingBench and 15.4 on OVO-Bench, peaking at 18.7
Full article253 words · extracted from huggingface.co · click to collapse
Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability to dynamically configure future computation in a training-free manner. Realizing such anticipatory computation, however, is very challenging: future anticipation must be sufficiently reliable to guide computation, planning must run concurrently with streaming inference, and online reconfiguration must incur negligible overhead. To address these challenges, we introduce FORESIGHT, a dual-stream architecture comprising two Siamese LLMs with shared weights, input encoders, and KV cache. The first LLM continuously processes incoming tokens, while the second runs ahead of the stream to anticipate future context, plan future computation, and generate task responses without interrupting streaming inference. Each plan decides when to reason next, what to check then, and how densely to sample, keeping transient evidence separate from persistent control. The resulting computation plan is executed online through an efficient reconfiguration protocol with schemaguided decoding and lightweight diff-based updates, enabling dynamic adaptation with low overhead. With a frozen Qwen3-VL-8B backbone, FORESIGHT achieves 23.0 mean joint F1 on OmniPro Online evaluation beating strongest trained baseline by 9.5%, while improving the backbone by 6.7 on StreamingBench and 15.4 on OVO-Bench, with the largest gain of 18.7 when evidence arrives later in the video stream. Our source code will be made publicly available.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.03123