Probe-Space Preconditioning for Fast and Stable Zero-Order Training
1.5-SPSA preconditions zero-order probes so OPT-13B beats MeZO on SST-2 in 70 steps instead of 100,000.
The paper targets the convergence gap of zero-order optimization, which can train OPT-30B in inference mode at roughly 60GB instead of about 600GB for Adam with backpropagation. Reallocating compute toward larger probe batches lets 1SPSA beat MeZO with less training compute, and 1.5-SPSA adds one clean forward pass per step to form a cheap diagonal preconditioner. Across six post-training datasets on Qwen3 and OPT, 1.5-SPSA leads prior zero-order solvers; on OPT-13B SST-2 it gains 3.1 accuracy points over both MeZO and backpropagation in 70 steps versus MeZO's 100,000, and the implementation trains OPT-30B in place on A100 GPUs.
- Adam training of OPT-30B is cited at about 600GB, versus about 60GB for zero-order inference-mode training.
- 1.5-SPSA adds one clean forward pass to build a diagonal probe-space preconditioner.
- On OPT-13B SST-2 it gains 3.1 points over MeZO and backprop in 70 steps versus MeZO's 100,000.
- Packed probes and fused kernels train up to OPT-30B in place on A100 GPUs.
Full article239 words · extracted from arxiv.org · click to collapse
Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU memory (assuming batch size 8 and sequence length 2048). Alternatively, zero-order optimization (ZOO) trains in inference-mode (requiring only $\approx$ 60GB for the same model): no stored activations, no gradients, and no optimizer states. However, ZOO convergence has lagged behind BP. In this work, we evaluate two methods to close this gap. First, we show that reallocating training compute budget from many steps to large effective batch sizes with many perturbations (or probes) but fewer steps, allows 1SPSA (Spall, 1992) to outperform zero order methods like MeZO (Malladi et al., 2023) with less training compute. Next, we introduce 1.5-SPSA, adding a single "clean" forward-pass per step to 1SPSA to calculate a cheap diagonal preconditioner in probe-space, which improves convergence rate and convergence by down-weighting high curvature directions. Benchmarking on 6 post-training datasets on both Qwen3 and OPT model families, we show that 1.5-SPSA achieves State-of-the-Art results over previous ZOO solvers with much less optimization steps. For example, we train OPT-13B (for direct comparison to MeZO) and find 1.5-SPSA achieves +3.1% accuracy on SST-2 over both MeZO and BP in only 70 steps vs. MeZO's 100,000 steps. Finally, we combine an 8-bit-packing random generator, triton fused unpack/apply kernels, and distributed parallelism to achieve fast and stable training of models as large as OPT-30B in-place on commodity GPUs (e.g. A100).
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.38095