Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation
Learning from self-generated text destabilizes test-time training across 128K tokens, including Adam updates on Qwen3-4B.
When test-time training learns from a model's own output, each update changes the model that writes the next example. Across 128K-token streams, retaining generated-text updates worsened prediction on independent human text for TTT-E2E configurations labeled 125M, 760M, and 3B, and the same failure occurred when Adam updated Qwen3-4B. Generating chunks with a frozen model removed over 98% of the damage at 125M and 760M. A Settlement check on independent real text before commitment left mean endpoint gaps of 0.07 and -0.02 nats while preserving adaptation on real text.
- Keeping generated-text updates hurt prediction on independent human text.
- Seen in TTT-E2E 125M, 760M, and 3B, plus Adam on Qwen3-4B.
- A frozen generator removed over 98% of damage at 125M and 760M.
- Settlement checks independent real text before committing an update.
Full article207 words · extracted from huggingface.co · click to collapse
Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-written text with three TTT-E2E model configurations (labeled 125M, 760M, and 3B). The same failure occurs when Adam updates Qwen3-4B's existing weights. The same update mechanisms can improve on real text, so writing itself is not the failure. Three matched comparisons trace the causal pathway. Fixed Generation removes over 98% of the damage at 125M and 760M by using a frozen model to generate training chunks. Recorded Replay separates the loss caused by reading degraded text from the additional loss stored by updating on it. A paired one-update comparison then shows the local conflict: an update predicts its source better but new real text worse. This cost grows after Closed Loop adaptation, with a few trajectories accounting for most large failures. Finally, Settlement evaluates the candidate state on independent real text before commitment. It leaves mean endpoint gaps of 0.07 and -0.02 nats at 125M and 760M while retaining real-text adaptation. These results motivate checking prediction on independent evidence before retaining an update.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.05076