Expert-Space Exploration in MoE Reinforcement Learning
ESRL explores MoE expert-routing space during RL, lifting Qwen3-30B-A3B Pass@1 by 3.2 points over GRPO at no extra cost.
ESRL (Expert-Space Exploration Reinforcement Learning) is an architecture-aware framework that treats expert routing in Mixture-of-Experts models as an additional source of rollout diversity. It preserves high-confidence experts as anchors, restricts stochastic routing to a plausible candidate pool, adapts perturbation strength to router entropy, and replays recorded expert paths during policy optimization. ESRL achieves the best performance across top-K, top-1, and shared-expert routing backbones on mathematics, science, and code tasks; on Qwen3-30B-A3B it improves average Pass@1 and Pass@8 over GRPO by 3.2 and 4.5 percentage points respectively, without additional sampling or compute.