RLHF & Post-Training Course by Nathan Lambert
Course
A full course accompanying the book with added resources and other lectures I've given.
The slide decks are usually built with Claude Opus, with substantial human revisions. (Exceptions are noted on the deck itself.)
Welcome to the Course
Introduction and overview of what you'll learn
Watch
Prerequisites — what to know before starting.
Primary Material — the core lecture series that follows the book chapter by chapter, with recordings, slides, PDFs, and source.
Extra Resources — recommended books, external RL courses, and Nathan's own talks paired with the chapters they go with.
Other Guest Lectures and Talks — invited talks and standalone presentations.
Prerequisites
This course is roughly aimed at early AI PhD or master's students, but it is designed to be accessible to anyone willing to put in the work.
You do not need prior reinforcement learning or language modeling background to start. A motivated learner who studies hard — leaning on today's top LLMs as a tutor to unpack unfamiliar math, code, and jargon — can follow the entire course. I encourage you to go down rabbit holes, skip or reorder videos, and chase what excites you.
If you prefer the traditional coursework path, the usual background is the basics of language modeling / NLP plus basic machine learning (e.g. an intro to AI course and an intro to ML course).
The ML Foundations of LLM Post-Training
A refresher on the ML prerequisites of post-training — language modeling, KL, cross-entropy, & other math — to acclimate to the series
Watch<br>PDF<br>Slides<br>Source
Additional learning material
Primary Material
Lecture 1: Overview
Chapters 1-3 · Foundations of RLHF and post-training
Watch<br>PDF<br>Slides<br>Source
Lecture 2: IFT, Reward Models, & Rejection Sampling
Chapters 4, 5, 9 · Start of the core optimization methods section
Watch<br>PDF<br>Slides<br>Source
Lecture 3: RL Motivation & Math
Chapter 6, Part 1 · Policy gradients math, intuitions, and theory
Watch<br>PDF<br>Slides<br>Source
Lecture 4: RL Implementation & Practice
Chapter 6, Part 2 · Code, loss aggregation, async training, and practical engineering
Watch<br>PDF<br>Slides<br>Source
Q&A 1: Reader questions, lec. 1-4
Watch<br>PDF<br>Slides<br>Source
Lecture 5: The Rise of Reasoning Models
Chapter 7 · RLVR, inference-time scaling, and the 2025 reasoning model wave
Watch<br>PDF<br>Slides<br>Source
Lecture 6: Direct Preference Optimization
Chapter 8 · Deriving DPO step by step, plus variants and practice
Watch<br>PDF<br>Slides<br>Source
Conversation 1: Frontier post-training recipes in 2026 (w/ Finbarr Timbers)
Watch<br>PDF<br>Slides<br>Source
Lecture 7: Synthetic Data and Modern Post-training Methods
Chapter 12 · On-policy distillation, AI feedback, Constitutional AI, and rubrics
Watch<br>PDF<br>Slides<br>Source
Q&A 2: Reader questions, lec. 5-7
Watch<br>PDF<br>Slides<br>Source
Lecture 8: On "Preferences" and Preference Data
Chapters 10 & 11 · A history of preferences, the preference-data engine, and unanswerable questions
Watch<br>PDF<br>Slides<br>Source
Extra Resources
Outside material that I've personally used for going deeper on reinforcement learning and language models. The books in particular are wonderful complements.
Books & Courses
Book<br>Reinforcement Learning: An Introduction<br>Sutton & Barto · The foundational RL textbook
Book<br>Build a Large Language Model (From Scratch)<br>Sebastian Raschka
Book<br>Build a Reasoning Model (From Scratch)<br>Sebastian Raschka
Course<br>CS285: Deep Reinforcement Learning (2023)<br>UC Berkeley · Sergey Levine
Course<br>Introduction to Reinforcement Learning (2015)<br>DeepMind / UCL · David Silver
More talks from Nathan
Various presentations over my last few years working in post-training, grouped by corresponding chapter.
Chapter 1 · Introduction
Aligning Open Language Models (Stanford CS25) Apr 2024
Chapter 3 · Training Overview
How Language Model Post-Training Is Done Today Jan 2025
How We Built a Leading Reasoning Model (Olmo 3) Dec 2025
Chapter 6 · Reinforcement Learning
GRPO's New Variants and Implementation Secrets Mar 2025
Chapter 10 · The Nature of Preferences
15 Minute History of Reinforcement Learning and Human Feedback Dec 2023
Chapter 7 · Reasoning & Inference-Time Scaling
An Unexpected Reinforcement Learning Renaissance Feb 2025
Early Stages of the Reinforcement Learning Era of Language Models Mar 2025
Experimenting with Reinforcement Learning with Verifiable Rewards (RLVR) Apr 2025
Chapter 8 · Direct-Alignment Algorithms
DPO Debate: Is RL Needed for RLHF? Dec 2023
Life After DPO (Stanford CS224N) May 2024
An Update on DPO vs PPO for LLM Alignment Jul 2024
Other Guest Lectures and Talks
2026
An Introduction to Reinforcement Learning from Human Feedback and Post-training
SALA 2026 · Quito, Ecuador · March 2026
Invited Talk<br>PDF<br>Full Screen<br>Source
Citation
If you found this useful for your...