ZeroHour

Search: “Llama”

2 stories in the last 24h

A Zeroth-Order Paradigm for LLM Preference Alignment

Researchers propose ComPO, a zeroth-order comparison-based preference alignment method with convergence guarantees that mitigates likelihood displacement in LLMs.

ComPO extracts directional information from preference pairs via comparison oracles instead of optimizing a differentiable preference loss, addressing likelihood displacement in direct alignment methods. The paper establishes convergence guarantees for the offline scheme and introduces an online variant with reverse-KL control using unlabeled policy generations. Experiments on Mistral, Llama, Gemma-2, Gemma-3, and Qwen3 models show improvements over existing direct alignment methods, including length-controlled win rates.

arXiv cs.AI / cs.LG / cs.CL · 12h agoAI research

Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models

A new site tracks release dates and training cutoffs for 20 AI models across 8 labs, exposing months-long staleness gaps.

A community-built page, launched via Show HN, tracks release dates and training cutoff dates for 20 current models from 8 labs including OpenAI, Anthropic, Google DeepMind, Meta, Mistral AI, Alibaba, DeepSeek, and xAI. Only 10 of the 20 models have lab-published cutoff dates, with data available as models.json. For example, GPT-6 Astra shipped September 3, 2026 with an April 30, 2026 cutoff. The page argues web search tools paper over, but never close, the staleness gap.

Hacker News · AIupdated · 13h agofirst · 17h agoAI tools & infra 19 sourcesHN 41↑ · 30 comments