Latent Space·2d agoJev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI#jev#typesafe#rlcd 4 sources 15 min
Hugging Face daily papers·3d agoRufus-Air: An Open LLM Post-Training Recipe#post-training#sft#reinforcement-learning
TechCrunch · AI·8d agoA new kind of AI model from a ChatGPT inventor is thrilling developers#calibrated-decisions#jev#llm-alternative 4 min
arXiv cs.AI / cs.LG / cs.CL·10d agoA Zeroth-Order Paradigm for LLM Preference Alignment#compo#likelihood-displacement#llmAI research
Hugging Face daily papers·11d agoRethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening#critic-learning#ppo#preference-optimization
Hugging Face daily papers·11d agoA Zeroth-Order Paradigm for LLM Preference Alignment#likelihood-displacement#post-training#preference-alignment1
404 Media·12d agoInside ‘Project Lily’: The Humans Reading Your ChatGPT Chats#anthropic#chatgpt#human-review 11 min
TechCrunch · AI·17d agoOpenAI adds a prominent AI doomer to its board of directors#ai-safety#alignment#board-appointment 3 min
Hugging Face daily papers·27d agoPLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization#alignment#dpo#noisy-labels
Hugging Face daily papers·27d agoAgenticGen: Reward-Guided Agentic Video Generation for Advertising#advertising#agents#dpo
Interconnects·Aug 12, 2026I wrote an AI textbook — how long until AI can do it better?#llm#long-form#model-capabilities 11 min
Interconnects·Aug 10, 20265 useful things you'll learn in my new post-training textbook (shipping now!)#grpo#llm#post-training 7 min1