PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning
PEARL adapts prefill-decode execution and elastic GPUs, raising agentic RL throughput by up to 2.79x.
PEARL is an asynchronous agentic reinforcement-learning system that combines external GPU elasticity, temporary reuse of idle training GPUs, and adaptive prefill-decode execution. Runtime profiles predict rollout completion, including environment-driven drops in decode concurrency, and cost-aware policies limit low-benefit reconfiguration. It achieves 2.17-2.79x the throughput of fixed-resource ROLL and improves over RLBoost+ by up to about 26.9% for Qwen3-8B and 36.3% for Qwen3-30B-A3B.
51