A Unified and Constrained View of Regularization-Based Robust Reinforcement Learning
Paper unifies regularization-based robust RL methods via new performance-gap upper bounds and jointly learned Lagrange multipliers.
The authors derive new upper bounds on the gap between nominal and worst-case deep RL policies, each expressible as an existing regularization objective plus a KL-divergence penalty. Robust training is reformulated as constrained optimization, where prior methods correspond to a fixed Lagrange multiplier. The multiplier is instead updated jointly with the policy, auto-tuning the regularization weight. Adversarial evaluations across several continuous control tasks validate the theory.