QAM: Quadratic-Accurate Checkpoint Merging via Sequential Consistency
QAM merges checkpoints to match sequential gradient descent through second order, with an O(h^3) error bound.
Quadratic-Accurate Merging studies how saved checkpoints can reconstruct a sequential training endpoint. Under a local transition model, two checkpoint-index moment conditions characterize convex merges that agree through second order, and an information limit shows no fixed-length gradient-descent history can achieve uniform o(h^3) endpoint error. QAM attains a matching O(h^3) bound and is the unique profile-dependent merge that exactly matches sequential GD on fixed quadratics. On Adam trajectories from SmolLM3-3B and OpenEuroLLM-Prelude-9B, results were mixed on short windows and stronger than Warmup-Stable and Merge on longer windows.
- Moment conditions characterize convex merges accurate through second order.
- No fixed-length GD history can beat o(h^3) endpoint error uniformly.
- QAM matches that O(h^3) bound and is unique on quadratics.
- Mixed gains on SmolLM3-3B and OpenEuroLLM-Prelude-9B versus WSM.
Full article229 words · extracted from arxiv.org · click to collapse
Saved checkpoints record states along a training trajectory, but generally do not determine the updates at states that would be visited under a different schedule. We study how accurately these checkpoints can reconstruct the endpoint of a sequential reference with prescribed update strengths. Under a common local transition model, two checkpoint-index moment conditions characterize all convex merges that agree with this reference through second order. We then prove an information limit that for nondegenerate profiles, no algorithm using only a fixed-length gradient-descent (GD) history with step size $h$ can achieve $o(h^3)$ endpoint error uniformly over a fixed class of smooth, strongly convex losses. The lower bound follows from two losses with identical GD checkpoint histories but sequential reference endpoints separated by $Ω(h^3)$. \textbf{Quadratic-Accurate Merging} (QAM) achieves a matching uniform $O(h^3)$ endpoint error bound. Its explicit coefficients also define the unique profile-dependent merge that exactly matches the sequential GD reference across all fixed quadratic objectives. Across two public Adam checkpoint trajectories (SmolLM3-3B and OpenEuroLLM-Prelude-9B), three windows and three profiles per model, and 15 tasks, QAM shows mixed results for short windows and broader advantages over \textbf{Warmup-Stable and Merge} (WSM) for longer windows. Matched-moment GSM8K diagnostics further show that local consistency alone does not fully determine downstream scores. These results characterize the reconstruction limits of saved histories, provide a coefficient rule that attains the optimal rate, and assess its practical utility.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35168