MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning
MetaRubric fixes 'Vacuous Credit' in rubric-based RL using counterfactual prompts and evidence-aware scoring, lifting PubMedQA up to 20.4 points.
The paper identifies 'Vacuous Credit', where rubric judges assign high criterion scores even when required information or actions are absent, persisting after removal and potentially reversing GRPO advantage signs. MetaRubric alternates evidence-aware policy optimization with response-guided rubric adaptation, constructing counterfactual prompts by changing one task-relevant fact and requiring sufficient evidence before assigning credit. Across multiple backbones it improves PubMedQA accuracy by 6.00-20.40 percentage points over static-judge GRPO, with further gains on HealthBench-Hard and two multimodal medical benchmarks.
- Vacuous Credit lets rubric judges reward responses missing required information
- Counterfactual prompts flip one task-relevant fact to test credit assignment
- Evidence-aware scoring plus response-guided rubric revision across training stages
- PubMedQA gains 6.00-20.40 percentage points over static-judge GRPO
Full article180 words · extracted from huggingface.co · click to collapse
Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the response, a failure mode we term Vacuous Credit. Such awards persist after the required information is removed and can reverse the sign of a response's GRPO advantage. To address this problem, we introduce MetaRubric, which alternates evidence-aware policy optimization with response-guided rubric adaptation. We construct counterfactual counterparts by changing one task-relevant fact in each prompt. During policy optimization, credit is assigned only when the response contains sufficient evidence to satisfy the required rubric criterion. After each policy-optimization stage, current policy responses guide revisions to original and counterfactual criteria while preserving the meaning of the original prompt's initial rubric as interpreted under each prompt's facts. We also adapt criterion weights at stage boundaries to better address observed policy errors. Across multiple backbones, MetaRubric improves PubMedQA accuracy by 6.00--20.40 percentage points over static-judge GRPO, with further gains on HealthBench-Hard and two multimodal medical benchmarks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.02824