ZeroHour
Product

ComPO

1 mentions in 7 days · 1 in 30 days · 1 total · first seen · last

Timeline

A Zeroth-Order Paradigm for LLM Preference Alignment

ComPO is a zeroth-order preference alignment method using comparison oracles to mitigate likelihood displacement across Mistral, Llama, Gemma, and Qwen3 models.

The paper proposes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method that extracts directional information from preference pairs with small likelihood margins without directly optimizing a differentiable preference loss. The authors prove convergence guarantees for the offline scheme and performance guarantees for a constrained online variant with reverse-KL control. Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 show improvements over existing direct alignment methods, including length-controlled win rates, with pair-level diagnostics consistent with mitigating likelihood displacement.

Hugging Face daily papersupdated · 17h agofirst · 1d agoAI research 2 sources

Appears with

Entities are extracted by the model from each article. Watching an entity keeps it in this browser only (no account); the watchlist page and dashboard alerts use it.