Inner Momentum for Differentially Private Muon
Averaging per-example Muon gradients over recent model snapshots before clipping bounds distortion and improves differentially private GPT-2 fine-tuning.
The paper proposes DP-Muon-IM, which averages each sampled example's Muon gradient over the current model and a short history of recent models before clipping, curbing clipping-induced distortion of singular-vector geometry. The clipped batch matrix separates into a common rescaling plus a covariance residual with Frobenius norm bounded by sigma_lambda * sigma_G, and finite Newton-Schulz iterations are shown to preserve the polar factor under these spectral conditions. On private GPT-2 fine-tuning over E2E and DART at epsilon in {1, 2, 4, 8}, DP-Muon-IM improves BLEU and ROUGE-L over DP-Muon in every seed-matched comparison, with 2-4% lower pre-noise polar error in non-private diagnostics.
- Per-example clipping distorts the singular-vector geometry Muon's polar-factor update depends on
- Averaging gradients over recent models bounds clipping distortion via a covariance residual
- Newton-Schulz orthogonalization provably preserves the correction
- BLEU/ROUGE-L gains over DP-Muon at epsilon 1-8 on E2E and DART
Full article173 words · extracted from arxiv.org · click to collapse
Differentially private training clips each per-example gradient before adding noise. This clipping is radial for each example, yet unequal clipping factors can distort the relative singular-vector geometry of their average. Muon is particularly exposed to this effect, since its update is an approximate polar factor UV^T that depends only on the singular vectors that clipping can shift. To curb this degradation, we propose averaging each sampled example's Muon gradient over the current model and a short history of recent models before clipping. The clipped batch matrix then separates into a common rescaling and a covariance residual R between sampled gradients and clipping values, with ||R||_F <= sigma_lambda sigma_G, bounding the clipping-induced distortion directly. We further show that a finite Newton-Schulz iteration preserves the polar factor of its input under these spectral conditions, confirming that our correction survives orthogonalization. In private GPT-2 fine-tuning on E2E and DART at epsilon in {1, 2, 4, 8}, DP-Muon-IM improves BLEU and ROUGE-L over DP-Muon in every seed-matched comparison, and non-private diagnostics show 2-4% lower pre-noise polar error.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.02738