Linear Recurrent Memory Suffices to Distil a World-Model Policy for Robot Air Hockey
A linear 64-dimensional recurrent memory matches DreamerV3 and GRU on simulated air-hockey defence under tracking loss.
Researchers study simulated air-hockey defence when puck tracking is temporarily lost. A DreamerV3 teacher beats a memoryless policy, and resetting its recurrent state sharply reduces performance, showing the task needs memory. Distilling it into a 64-dimensional diagonal linear recurrence matches both a GRU baseline and the teacher across five seeds, while added nonlinear innovation rank gives no benefit. The linear model uses fewer recurrent parameters than GRU; conclusions are limited to this simulation, teacher, state size, and blackout horizon.
- DreamerV3 teacher needs memory under temporary puck-tracking loss
- Purely linear 64-d recurrence matches GRU and the teacher
- Higher nonlinear innovation rank showed no measured benefit
- Findings limited to this simulation, teacher, and blackout horizon
Full article206 words · extracted from arxiv.org · click to collapse
Does memory-dependent control need nonlinear recurrent dynamics? We study simulated air-hockey defence under temporary loss of puck tracking. A DreamerV3 teacher outperforms a memoryless policy under tracking loss, while resetting the teacher's recurrent state sharply reduces performance, which demonstrates that the task requires memory. We distil this teacher into compact recurrent policies with a 64 dimensional state, with a combination of a diagonal linear recurrence and an optional rank-$k$ nonlinear innovation while retaining nonlinear observation encoders and action heads. Across five matched seeds, the purely linear recurrent model ($k=0$) matches both the GRU baseline and the teacher throughout the tested range of tracking loss. Increasing nonlinear innovation rank providing no measured benefits. This result is obtained on a fresh test split, which will be only opened after all models and analyses are frozen. The linear model requires fewer recurrent parameters and less computation than GRU, but performs comparably. These results suggest that, for this memory dependent control task, nonlinear representation learning around a simple linear memory mechanism can be sufficient, and that nonlinear recurrent dynamics are not necessarily required. These conclusions are limited to the simulated task, teacher, state dimension, and blackout horizon considered here, and to policies whose observation encoder and action head remain nonlinear.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.39151