arXiv cs.AI / cs.LG / cs.CL·10d agoA Zeroth-Order Paradigm for LLM Preference Alignment#compo#likelihood-displacement#llmAI research