Pivotal role of reinforcement learning in modern large language models — Harvard Gazette
Menu
Sections
Featured Topics
Featured series
Wondering
A series of random questions answered by Harvard experts.
Explore the Gazette
Read the latest
When AI goes rogue
A bond stronger than blood
Why we need to break the grip of ‘oldigarchy’
Search
Search the Harvard GazetteGo
The explosion of modern AI, exemplified by the unprecedented abilities of large language models (LLMs), was enabled by a family of computational techniques known as machine learning (ML). But how exactly does a machine learn anything? In this explainer, we dive into one of the most important tools in the ML toolbox: reinforcement learning (RL).
The prehistory of RL: B.F. Skinner and operant conditioning
In the middle decades of the 20th century, B.F. Skinner was one of the most influential experimental psychologists in the world. A Harvard professor, Skinner is widely considered the father of operant conditioning, a technique for training an animal or a human to produce a specific behavior using rewards or punishments. It’s the same approach a dog owner might use to train a dog by rewarding certain behaviors with treats.
While still a graduate student, Skinner invented the operant conditioning chamber or “Skinner box,” a highly controllable setting in which a lab animal’s behavior can be observed and manipulated. A rat in a Skinner box might be given a treat whenever it happens to press a lever after a light is turned on. Over time, the rat learns to tap the lever as soon as the light turns on. Alternatively, the rat might be punished for tapping on a lever. Over time, the rat learns to avoid that action. In both cases, the rat learns what to do, and what not to do, by trial and error. Over time, the rat’s voluntary behavior is shaped by the consequences of its past actions.
A rat running around in a cage tapping on levers might seem a world away from a large language model running on a state-of-the-art supercomputer like the Kempner AI cluster, but there’s an important thread linking Skinner’s work with the ongoing AI renaissance.
Read Full Story
Share this article
Share on Facebook
Share on LinkedIn
Email article
Print/PDF
You might like
Nation & World
When AI goes rogue
Computer security expert says recent OpenAI, Anthropic breaches highlight need for regulations that balance safety, speed of development
9 min read
Arts & Culture
A bond stronger than blood
Tayari Jones digs into themes of latest novel, ‘Kin’: ‘We all need chosen family.’
3 min read
Nation & World
Why we need to break the grip of ‘oldigarchy’
Legal historian’s new book argues older Americans control power, wealth — and lack sense of urgency to fix nation’s most pressing problems
Part of the<br>Excerpts<br>series
long read
Trending
Work & Economy
Are we headed toward recession? Unpredictable.
Economic historian rebuts various beliefs about cause, concerns around downturns, explains why expansions are more important anyway
9 min read
Health
The best medicine may be free
Forest therapy guide explains what a daily dose of nature does for the body
6 min read
Health
Coffee refill? Go for it, says American Heart Association, with note of caution.
Co-author of new guidance explains how many cups a day are safe — possibly even good for you — and why energy drinks are different
3 min read
Explore the Gazette
Our recent series
Wondering
A series of random questions answered by Harvard experts.
Life | Work
A series focused on the personal side of Harvard research and teaching.
Follow us on
TikTok
YouTube