Skip to content
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining · ZeroHour