ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
ZooWork-ShopRanker releases 0.6B, 4B, and 8B e-commerce rerankers trained on LLM shopping preferences.
Researchers present ZooWork-ShopRanker, open e-commerce rerankers at 0.6B, 4B, and 8B parameters aligned to shopping preferences rather than topical relevance alone. Training pairs are labeled by a panel of reasoning LLMs using position-debiased judgments, and the 8B model distills into the smaller models. ShopRank-Bench contains about 10,000 private-traffic preference pairs. The 8B and 4B models outperform the strongest open reranker baseline, with gains also reported on common MTEB benchmarks.
- Family includes 0.6B, 4B, and 8B rerankers.
- LLM judges label pairs with agreement tiers and position debiasing.
- The 8B teacher distills scores into the 4B and 0.6B models.
- ShopRank-Bench contains about 10,000 private preference pairs.
- The 8B and 4B beat the strongest open reranker baseline.
Full article209 words · extracted from huggingface.co · click to collapse
Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and comparative product fit. These preference signals are difficult to supervise at scale: real search traffic provides authentic queries and candidates but no clean pairwise labels. We present ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, and 8B) aligned to judge-labeled shopping preference. Training pairs are labeled by a panel of reasoning large language models (LLMs) from different families acting as a preference oracle, with position-debiased judgments and agreement tiers, and the rerankers are trained on these labels. The aligned 8B flagship then serves as a distillation teacher for the efficient 4B and 0.6B models, which are fit to its scores and sharpened on judged pairs. To measure progress, we introduce ShopRank-Bench, a contamination-limited benchmark of ~10,000 private-traffic preference pairs in both text formats, tiered by how many judge families committed to each label. ZooWork-ShopRanker-8B and -4B significantly outperform the strongest open reranker baseline, every model significantly beats its own un-aligned base, and ZooWork-ShopRanker-0.6B beats its size peer; the gains hold in both formats and extend to common MTEB benchmarks. We release the models and the dual-format ShopRank-Bench to facilitate further research.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.31002