RLHF and Post-Training Course by Nathan Lambert

ankitg121 pts0 comments

RLHF & Post-Training Course by Nathan Lambert

Course

A full course accompanying the book with added resources and other lectures I've given.

The slide decks are usually built with Claude Opus, with substantial human revisions. (Exceptions are noted on the deck itself.)

Welcome to the Course

Introduction and overview of what you'll learn

Watch

Prerequisites — what to know before starting.

Primary Material — the core lecture series that follows the book chapter by chapter, with recordings, slides, PDFs, and source.

Extra Resources — recommended books, external RL courses, and Nathan's own talks paired with the chapters they go with.

Other Guest Lectures and Talks — invited talks and standalone presentations.

Prerequisites

This course is roughly aimed at early AI PhD or master's students, but it is designed to be accessible to anyone willing to put in the work.

You do not need prior reinforcement learning or language modeling background to start. A motivated learner who studies hard — leaning on today's top LLMs as a tutor to unpack unfamiliar math, code, and jargon — can follow the entire course. I encourage you to go down rabbit holes, skip or reorder videos, and chase what excites you.

If you prefer the traditional coursework path, the usual background is the basics of language modeling / NLP plus basic machine learning (e.g. an intro to AI course and an intro to ML course).

The ML Foundations of LLM Post-Training

A refresher on the ML prerequisites of post-training — language modeling, KL, cross-entropy, & other math — to acclimate to the series

Watch<br>PDF<br>Slides<br>Source

Additional learning material

Primary Material

Lecture 1: Overview

Chapters 1-3 &middot; Foundations of RLHF and post-training

Watch<br>PDF<br>Slides<br>Source

Lecture 2: IFT, Reward Models, & Rejection Sampling

Chapters 4, 5, 9 &middot; Start of the core optimization methods section

Watch<br>PDF<br>Slides<br>Source

Lecture 3: RL Motivation & Math

Chapter 6, Part 1 &middot; Policy gradients math, intuitions, and theory

Watch<br>PDF<br>Slides<br>Source

Lecture 4: RL Implementation & Practice

Chapter 6, Part 2 &middot; Code, loss aggregation, async training, and practical engineering

Watch<br>PDF<br>Slides<br>Source

Q&A 1: Reader questions, lec. 1-4

Watch<br>PDF<br>Slides<br>Source

Lecture 5: The Rise of Reasoning Models

Chapter 7 &middot; RLVR, inference-time scaling, and the 2025 reasoning model wave

Watch<br>PDF<br>Slides<br>Source

Lecture 6: Direct Preference Optimization

Chapter 8 &middot; Deriving DPO step by step, plus variants and practice

Watch<br>PDF<br>Slides<br>Source

Conversation 1: Frontier post-training recipes in 2026 (w/ Finbarr Timbers)

Watch<br>PDF<br>Slides<br>Source

Lecture 7: Synthetic Data and Modern Post-training Methods

Chapter 12 &middot; On-policy distillation, AI feedback, Constitutional AI, and rubrics

Watch<br>PDF<br>Slides<br>Source

Q&A 2: Reader questions, lec. 5-7

Watch<br>PDF<br>Slides<br>Source

Lecture 8: On "Preferences" and Preference Data

Chapters 10 & 11 &middot; A history of preferences, the preference-data engine, and unanswerable questions

Watch<br>PDF<br>Slides<br>Source

Extra Resources

Outside material that I've personally used for going deeper on reinforcement learning and language models. The books in particular are wonderful complements.

Books & Courses

Book<br>Reinforcement Learning: An Introduction<br>Sutton & Barto &middot; The foundational RL textbook

Book<br>Build a Large Language Model (From Scratch)<br>Sebastian Raschka

Book<br>Build a Reasoning Model (From Scratch)<br>Sebastian Raschka

Course<br>CS285: Deep Reinforcement Learning (2023)<br>UC Berkeley &middot; Sergey Levine

Course<br>Introduction to Reinforcement Learning (2015)<br>DeepMind / UCL &middot; David Silver

More talks from Nathan

Various presentations over my last few years working in post-training, grouped by corresponding chapter.

Chapter 1 &middot; Introduction

Aligning Open Language Models (Stanford CS25) Apr 2024

Chapter 3 &middot; Training Overview

How Language Model Post-Training Is Done Today Jan 2025

How We Built a Leading Reasoning Model (Olmo 3) Dec 2025

Chapter 6 &middot; Reinforcement Learning

GRPO's New Variants and Implementation Secrets Mar 2025

Chapter 10 &middot; The Nature of Preferences

15 Minute History of Reinforcement Learning and Human Feedback Dec 2023

Chapter 7 &middot; Reasoning & Inference-Time Scaling

An Unexpected Reinforcement Learning Renaissance Feb 2025

Early Stages of the Reinforcement Learning Era of Language Models Mar 2025

Experimenting with Reinforcement Learning with Verifiable Rewards (RLVR) Apr 2025

Chapter 8 &middot; Direct-Alignment Algorithms

DPO Debate: Is RL Needed for RLHF? Dec 2023

Life After DPO (Stanford CS224N) May 2024

An Update on DPO vs PPO for LLM Alignment Jul 2024

Other Guest Lectures and Talks

2026

An Introduction to Reinforcement Learning from Human Feedback and Post-training

SALA 2026 &middot; Quito, Ecuador &middot; March 2026

Invited Talk<br>PDF<br>Full Screen<br>Source

Citation

If you found this useful for your...

middot chapter source watch slides learning

Related Articles