SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs
SlideDP scales host-resident multi-GPU LLM fine-tuning, beating offload baselines and nearing FSDP2 on Qwen3-14B.
SlideDP is a synchronous data-parallel runtime that keeps one authoritative host state and pipelines parameter delivery, gradient aggregation, and CPU updates for full-parameter LLM fine-tuning beyond GPU memory. Matched-batch sweeps report geometric-mean throughput 1.46-2.64 times SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s it approaches GPU-resident FSDP2 for Qwen3-14B and, with a larger batch, processes over 1 million tokens per step, 11.2% above FSDP2's measured peak. It also supports 256K-token sequences for that model and fine-tunes Qwen2.5-72B on four RTX 4090s.