A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models
Study shows video models often learn correct physics but fail to use it; low-dimensional 'causal writability' edits can restore correct motion.
The paper demonstrates 'causal writability' in video generation models: physically correct motion remains available inside the model even when the model outputs incorrect motion. In a red/blue mass oscillation setup, a low-dimensional edit predicted from simple physical variables restores correct fast motion, with a sharp depth boundary marking commitment. Early causal writability predicts which training errors later get corrected, and both writability and closure reproduce in a pretrained 1.3B video model.
- Correct motion often learned but unused; edits can restore it
- Sharp depth boundary marks commitment point for each edit
- Early causal writability predicts errors that later training corrects
- Reproduced in pretrained 1.3B video model, supporting generality
Full article192 words · extracted from arxiv.org · click to collapse
When a video model generates physically incorrect motion, did it fail to learn the correct motion, or did it learn it but fail to use it? We show the latter: the correct motion remains available inside the model and can still be made to control the generated video. We train on videos where red masses oscillate slowly and blue masses oscillate quickly, then test a red mass with fast observed motion. Even when the model generates slow motion in this conflicting case, a low-dimensional edit predicted from simple physical variables restores the correct fast motion. We call this ability causal writability. At fixed strength, we find a sharp depth boundary: the same edit changes the video before the boundary but not after it. This closure marks commitment for that write. The motion signal nevertheless remains, and a stronger downstream write can restore physical motion, while excessive gain overshoots. Early causal writability predicts which errors training later corrects: those errors are writable at more network depths than errors that persist. We reproduce both causal writability and its sharp closure in a pretrained 1.3B video model, supporting generality across model scale and training regime.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.15980