Think Thrice Before Reranking: Multi-perspective Evidence and Reasoning Integration for Text Reranking
MERIT-Rank integrates multiple LLM reasoning trajectories for text reranking; its 4B model outperforms most 7B and 32B rerankers on BRIGHT.
MERIT-Rank formulates a Multi-Trajectory Reasoning Space that evaluates query-document relevance from multiple perspectives and consolidates reasoning paths with a joint reranker. A progressive training framework, Progressive Rank Policy Optimization (PRPO), stabilizes reasoning trajectories while improving ranking quality through staged objectives. Experiments on reasoning-intensive and traditional retrieval benchmarks show consistent gains over competitive baselines, with the 4B model outperforming most 7B and even 32B rerankers on BRIGHT.