TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
TT-VidT pairs a DINOv3 spatial ViT with a temporal layer, leading motion benchmarks at lower FLOPs.
TT-VidT separates a DINOv3-initialized ViT-B/16 spatial path from a compact Temporal Transfer Layer trained with Diff Compression to reconstruct target frames from a first-frame appearance anchor and motion tokens. A matched 24-configuration study at roughly 170–190M scale uses sim1.7M OpenVid and Moments-in-Time v2 clips for eight epochs. TT-VidT leads Jester, Something-Something V2, ARID, and Diving48, improving 54–121% over the strongest non-TT row while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2.