42
30
45
55
30
42
30
57
30
57
60
60
30
60
42
60
42
42
42
60
50
42
5 useful things you'll learn in my new post-training textbook (shipping now!)
Nathan Lambert's new RLHF and post-training LLM textbook covers PPO, GRPO, GSPO, CISPO and related techniques, freely available online.
Nathan Lambert's book 'Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs' is now shipping from Manning. It covers policy-gradient algorithms including PPO, GRPO, GSPO, CISPO, and RLOO, plus loss aggregation, truncated importance sampling, asynchronous RL systems, and post-training topics like rejection sampling, outcome reward models, and on-policy distillation. The book is freely available online with a 12-hour course, codebase, and exercises.
28
30
60
47
30
60
47
42
60
30
30
60
42
60
57
60
30
30