What masking geometry works best for EEG foundation models?
An EEG masking ablation matches REVE-level decoding performance at a fraction of the pre-training compute.
The paper isolates spatio-temporal masking choices for EEG foundation models, training 58 models in one pipeline under MAE and JEPA. Linear probes on the 12 OpenEEGBench datasets show both frameworks agree on an optimal mask and on shared failure modes, while 11 MAE and 9 JEPA configurations are statistically tied with the best. The authors also describe a JEPA-specific bias-inflation collapse that standard detectors miss. With a well-chosen mask, the pipeline reaches REVE-level downstream performance at a fraction of REVE's pre-training compute.
- Spatio-temporal masking is ablated under MAE and JEPA with one shared pipeline.
- Fifty-eight models are linearly probed on 12 OpenEEGBench datasets.
- Both frameworks agree on an optimal mask and shared failure modes.
- A well-chosen mask matches REVE downstream performance at much lower compute.
- The authors identify a JEPA-only bias-inflation collapse missed by standard detectors.
Full article177 words · extracted from huggingface.co · click to collapse
EEG foundation models hold promise for scalable brain-signal decoding across clinical and cognitive neuroscience applications, yet their pre-training pipelines remain poorly understood. Among design choices, the masking strategy is particularly critical: it determines what the network must predict and from which context. Yet it has never been ablated in isolation, as each new model bundles a new masking strategy with a new backbone and objective. In this paper, we formalize the design choices for spatio-temporal masking strategies and train various models with a single pipeline under varying masking configurations across two SSL frameworks (MAE and JEPA). We then systematically evaluate the resulting 58 pre-trained models on the 12 datasets of OpenEEGBench under a linear probe. Both frameworks agree on an optimal masking configuration and on shared failure modes. Outside these, performance is robust: 11 MAE and 9 JEPA configurations are statistically indistinguishable from the best. We further identify a novel JEPA-specific failure mode, tagged bias-inflation collapse, invisible to standard detectors. With a well-chosen mask, our pipeline reaches REVE-level downstream performance at a fraction of REVE's pre-training compute.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.33487