Probe-Space Preconditioning for Fast and Stable Zero-Order Training
1.5-SPSA preconditions zero-order probes so OPT-13B beats MeZO on SST-2 in 70 steps instead of 100,000.
The paper targets the convergence gap of zero-order optimization, which can train OPT-30B in inference mode at roughly 60GB instead of about 600GB for Adam with backpropagation. Reallocating compute toward larger probe batches lets 1SPSA beat MeZO with less training compute, and 1.5-SPSA adds one clean forward pass per step to form a cheap diagonal preconditioner. Across six post-training datasets on Qwen3 and OPT, 1.5-SPSA leads prior zero-order solvers; on OPT-13B SST-2 it gains 3.1 accuracy points over both MeZO and backpropagation in 70 steps versus MeZO's 100,000, and the implementation trains OPT-30B in place on A100 GPUs.