I Have a Stream: Making Self-Supervised Learning Work on Continuous Video
StreamMAE adapts masked autoencoders so pretraining works on continuous video streams without global shuffling.
The paper studies self-supervised pretraining on continuous video, where frames are consumed in temporal order with sliding-window batches and no global reshuffling or multi-epoch replay. Contrastive and distillation methods struggle, while MAE is more robust but still trails standard i.i.d. pretraining; the main issue is high intra-batch similarity from near-duplicate frames. StreamMAE keeps the MAE reconstruction objective and adds stream-aware regularization plus motion-biased crops. It beats streaming baselines, matches i.i.d. MAE on the same video, remains competitive with ImageNet-pretrained MAE, and improves as data grows from 12 to 95 hours.
- WT++ is a 95-hour urban walking-tour stream for ordered pretraining.
- Intra-batch near-duplicates, not inter-batch similarity, explain the gap.
- StreamMAE matches i.i.d. MAE on the same video and stays competitive with ImageNet MAE.
- Gains scale as the stream grows from 12 to 95 hours.
Full article184 words · extracted from huggingface.co · click to collapse
Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Combined with a comprehensive evaluation suite we find that contrastive and distillation-based methods struggle in this setting, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates. To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches i.i.d. MAE trained on the same video data, remains competitive with ImageNet-pretrained MAE, and scales positively as the pretraining stream grows from 12 to 95 hours.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.40333