SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs
SlideDP scales host-resident multi-GPU LLM fine-tuning, beating offload baselines and nearing FSDP2 on Qwen3-14B.
SlideDP is a synchronous data-parallel runtime that keeps one authoritative host state and pipelines parameter delivery, gradient aggregation, and CPU updates for full-parameter LLM fine-tuning beyond GPU memory. Matched-batch sweeps report geometric-mean throughput 1.46-2.64 times SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s it approaches GPU-resident FSDP2 for Qwen3-14B and, with a larger batch, processes over 1 million tokens per step, 11.2% above FSDP2's measured peak. It also supports 256K-token sequences for that model and fine-tunes Qwen2.5-72B on four RTX 4090s.
- Keeps one host state and pipelines parameters, gradients, and CPU updates
- Geometric-mean throughput is 1.46-2.64x over three offload baselines
- On four H100s, exceeds FSDP2 peak throughput by 11.2%
- Fine-tunes Qwen2.5-72B on four RTX 4090 GPUs
Full article161 words · extracted from huggingface.co · click to collapse
Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. Replicated transfers amplify traffic, while strong scaling can expose host work as computation windows shrink. We present SlideDP, a synchronous data-parallel runtime for shared-host multi-GPU systems. It maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. An analytical step-time model characterizes resource bottlenecks and pipeline exposure; runtime measurements guide communication, chunking, and activation policies under a GPU memory budget. In matched-batch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46-2.64times over SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2's measured peak throughput by 11.2%. Separately, it supports 256K-token sequences for the same model and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Project page: https://github.com/RegiaYoung/SlideDP.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.34162