Pivotal role of reinforcement learning in modern large language models

newsomix9xl1 pts0 comments

Pivotal role of reinforcement learning in modern large language models — Harvard Gazette

Menu

Sections

Featured Topics

Featured series

Wondering

A series of random questions answered by Harvard experts.

Explore the Gazette

Read the latest

When AI goes rogue

A bond stronger than blood

Why we need to break the grip of ‘oldigarchy’

Search

Search the Harvard GazetteGo

The explosion of modern AI, exemplified by the unprecedented abilities of large language models (LLMs), was enabled by a family of computational techniques known as machine learning (ML). But how exactly does a machine learn anything? In this explainer, we dive into one of the most important tools in the ML toolbox: reinforcement learning (RL).

The prehistory of RL: B.F. Skinner and operant conditioning

In the middle decades of the 20th century, B.F. Skinner was one of the most influential experimental psychologists in the world. A Harvard professor, Skinner is widely considered the father of operant conditioning, a technique for training an animal or a human to produce a specific behavior using rewards or punishments. It’s the same approach a dog owner might use to train a dog by rewarding certain behaviors with treats.

While still a graduate student, Skinner invented the operant conditioning chamber or “Skinner box,” a highly controllable setting in which a lab animal’s behavior can be observed and manipulated. A rat in a Skinner box might be given a treat whenever it happens to press a lever after a light is turned on. Over time, the rat learns to tap the lever as soon as the light turns on. Alternatively, the rat might be punished for tapping on a lever. Over time, the rat learns to avoid that action. In both cases, the rat learns what to do, and what not to do, by trial and error. Over time, the rat’s voluntary behavior is shaped by the consequences of its past actions.

A rat running around in a cage tapping on levers might seem a world away from a large language model running on a state-of-the-art supercomputer like the Kempner AI cluster, but there’s an important thread linking Skinner’s work with the ongoing AI renaissance.

Read Full Story

Share this article

Share on Facebook

Share on LinkedIn

Email article

Print/PDF

You might like

Nation & World

When AI goes rogue

Computer security expert says recent OpenAI, Anthropic breaches highlight need for regulations that balance safety, speed of development

9 min read

Arts & Culture

A bond stronger than blood

Tayari Jones digs into themes of latest novel, ‘Kin’: ‘We all need chosen family.’

3 min read

Nation & World

Why we need to break the grip of ‘oldigarchy’

Legal historian’s new book argues older Americans control power, wealth — and lack sense of urgency to fix nation’s most pressing problems

Part of the<br>Excerpts<br>series

long read

Trending

Work & Economy

Are we headed toward recession? Unpredictable.

Economic historian rebuts various beliefs about cause, concerns around downturns, explains why expansions are more important anyway

9 min read

Health

The best medicine may be free

Forest therapy guide explains what a daily dose of nature does for the body

6 min read

Health

Coffee refill? Go for it, says American Heart Association, with note of caution.

Co-author of new guidance explains how many cups a day are safe — possibly even good for you — and why energy drinks are different

3 min read

Explore the Gazette

Our recent series

Wondering

A series of random questions answered by Harvard experts.

Life | Work

A series focused on the personal side of Harvard research and teaching.

Follow us on

Instagram

LinkedIn

TikTok

Facebook

YouTube

Email

read skinner harvard series learning large

Related Articles