Andriy Burkov
THE HUNDRED-PAGE
DEEP REINFORCEMENT
LEARNING
BOOK

About the Book

The Hundred-Page Deep Reinforcement Learning Book

Deep reinforcement learning (deep RL) has transformed what machines can do. Self-driving cars carry passengers on public roads, robots learn to walk and recover from a stumble instead of replaying pre-programmed motions, and language models write working code and prove theorems long considered hard. None of this behavior was written by a programmer: the machine learned it all from experience. These methods are moving out of research labs into ordinary engineering work, making an understanding of their mathematical foundations and inner workings crucial for engineers who want to stay competitive as more systems are built this way.

This book traces the evolution of deep RL algorithms. It starts with on-policy learning and the Policy Gradient Theorem that gave rise to a family of actor-critic methods, including Proximal Policy Optimization (PPO) used to train both walking robots and large language models. Each algorithm builds on previous results, so every choice is justified. The book then turns to off-policy algorithms and derives the Deep Q-Network (DQN) from first principles. Every concept is motivated, grounded in clear mathematical foundations, and illustrated with graphs and working Python code.

Everything in the book is meant to be run. Each algorithm is implemented from scratch in PyTorch and trained on the same task—landing a reusable rocket booster on a moving ocean platform—so you can compare the algorithms side by side and see what each new idea changes. The closing chapter applies these methods to large language models: how a generated response becomes a trajectory, how RLHF trains reasoning models without a learned critic.

What's inside?

  • Mathematical foundations with intuitive explanations
  • Complete Python implementations with PyTorch on GitHub
  • Natural progression from REINFORCE to PPO and DQN
  • A perfect mix of theory, illustrations, and code
  • Extended versions of chapters on this website
  • $150 in free GPU credits on Lambda How?😲

Is the book for you?

Whether you're a technical leader, engineering manager, software developer, data scientist, or machine learning engineer, this book provides both the theoretical depth and the practical implementation skills essential for working with reinforcement learning—from robotics and control to finetuning language models.


What AI Leaders Say

Lambda Logo

Get $150 in Free GPU Credits

Purchase the book and receive $150 in free GPU credits on Lambda. Simply email your proof of purchase to author@thedrlbook.com to claim your credits.

Buy the Book

The Hundred-Page Deep Reinforcement Learning Book e-book
E-Book
Coming soon
The Hundred-Page Deep Reinforcement Learning Book hardcover
Hardcover
Coming soon
The Hundred-Page Deep Reinforcement Learning Book paperback
Paperback
Coming soon

Chapters

This book is published on the read-first, buy-later principle. All chapters will always remain available on this website.

About the Author

Andriy Burkov

Andriy Burkov is the author of "The Hundred-Page Machine Learning Book" and "The Hundred-Page Language Models Book," both of which became #1 Best Sellers on Amazon. He holds a Ph.D. in Artificial Intelligence and is a recognized expert in machine learning and natural language processing.

As a machine learning expert and leader, Andriy has successfully led dozens of production-grade AI projects in different business domains at Fujitsu, Gartner, and TalentNeuron. His books have been translated into more than a dozen languages and are used as textbooks in many universities worldwide. His work has impacted millions of machine learning practitioners and researchers worldwide.

Andriy's app: ChapterPal - an AI reading tutor for learners.

Stay in touch: LinkedIn, X, email, newsletter

Support Andriy by becoming a patron on: Patreon, Substack, or Buy him a Coffee