SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
SemanTok aligns video tokens with frozen DINO features so short prefixes carry clip semantics.
SemanTok is a flexible video tokenizer that feeds frozen DINO features into its encoder and reconstructs them from each retained token prefix with lightweight heads. A 201M-parameter SemanTok autoregressive model matches or beats a VideoFlexTok model 3.4 times larger, and larger SemanTok models further improve fidelity. Short prefixes are cheaper to predict and produce better generation quality, leaving pixel detail to later tokens. Semantic alignment holds on out-of-distribution classes and at every decoder noise level, including pure noise.
48