How Does Distribution Shift Shape Pretraining Gains in Neural PDE Surrogates?
Study shows pretraining gains for neural PDE surrogates depend on distribution shift type, target data budget, and whether source-target physics differ.
The authors pretrain a neural PDE surrogate on 254,909 RANS solutions from one airfoil family and fine-tune it on a new family under two target settings: same Spalart-Allmaras modeling and SA with added e^N transition modeling. At a 1000-sample budget, the pretrained model matches a from-scratch model trained on 3.25x as many samples for the same-physics target, but only 2.58x for the transition-modeled target, with orderings reversing at 5000 samples. Results indicate pretraining value depends jointly on target budget, coverage, and differences in modeled physics between source and target.
- Pretraining on 254,909 RANS airfoil solutions reduced data needs on a same-physics target
- Pretraining gains shrink when target physics changes (added e^N transition modeling)
- Benefit ordering between targets reverses as the fine-tuning budget grows to 5000 samples
- Sampling more distinct airfoils lowered error, but gains exceeded variance only for same-physics transfer
Full article163 words · extracted from arxiv.org · click to collapse
Pretraining a neural PDE surrogate can reduce the amount of new CFD data needed when geometry or modeled physics changes. However, it remains unclear how different components of distribution shift affect this benefit. We pretrain a surrogate on 254,909 RANS solutions from one airfoil family and fine-tune it on a new family under two target settings with matched freestream ranges: the same Spalart-Allmaras (SA) modeling and SA with added $e^N$ transition modeling. At $N=1000$, the pretrained model matches the accuracy of a model trained from scratch on $3.25\times$ as many samples for the same-SA target, but $2.58\times$ as many for the transition-modeled target. By $N=5000$, this ordering reverses ($1.56\times$ versus $1.86\times$). At $N=1000$, sampling more distinct airfoils lowers error on both targets, but only for the same-SA target is the gain increase larger than the observed draw-to-draw variation ($3.3\times$ to $4.0\times$). These results show that pretraining value depends jointly on target-data budget, target-data coverage, and whether source and target differ in modeled physics.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.20814