Forecasting from Counterfactual Simulator Rollouts: A Sim2Real Evaluation
Simulator rollouts of a new policy train forecasters that beat historical-data models in two real inventory deployments, cutting MAPE by up to 18.7 points.
The paper addresses cold-start forecasting when a newly deployed decision policy invalidates historical data, proposing to train predictors on counterfactual simulator rollouts of the target policy. It evaluates simulator fidelity, zero-shot Sim2Real transfer, and adaptation on two real-world inventory-control deployments. The simulator-trained forecaster reduces MAPE by 1.2-3.1 percentage points in Study 1 and 12.5-18.7 points in Study 2 versus the same architecture trained on historical real data. Lightweight calibration with early real observations after deployment further cuts error by up to 2.5 percentage points.
- Counterfactual simulator rollouts cover cold-start when new policy changes data distribution
- Evaluated on two real-world inventory-control deployments
- MAPE reduced 1.2-3.1 points (Study 1) and 12.5-18.7 points (Study 2)
- Post-deployment calibration with early real data cuts error by up to 2.5 points
Full article196 words · extracted from arxiv.org · click to collapse
Deploying a new decision policy creates a cold-start problem for prediction models whose targets depend on the policy's actions: historical observations reflect earlier policies, while real observations under the new policy are not yet available. Simulation offers a way to address this gap by rolling out the target policy across counterfactual scenarios and using the resulting trajectories to learn how the system responds to those controls. The simulation-to-reality (Sim2Real) transfer of this simulator-trained model can then be backtested by evaluating it against real observations from past deployments. Using two real-world inventory-control deployments, we evaluate this process from three angles: simulator fidelity, zero-shot transfer to real behavior, and adaptation as real target-policy observations accumulate. The simulator-trained forecaster achieves lower point-estimate mean absolute percentage error (MAPE) than the same architecture trained on historical real data, reducing MAPE by 1.2-3.1 percentage points in Study 1 and 12.5-18.7 points in Study 2. After deployment, lightweight calibration using early real observations further reduces error by up to 2.5 percentage points. These results provide empirical evidence that simulator-generated counterfactual data can support cold-start forecasting under a new policy, and the resulting model can be further refined as real deployment data become available.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.03662