Reinforcement Learning from Verifiable Rewards
Preface

Abstract
Reinforcement learning from verifiable rewards (RLVR) studies how models can improve by learning from reward signals derived from checkable task outcomes, executable feedback, formal validation, or other reliable forms of verification. This book’s purpose is to explain what kinds of rewards can be made verifiable, what those rewards train, where the paradigm has been most successful, and where it breaks.
LLM Use
Fortunately, we live in a world where AI slop writing is as intelligible as black from white. This remark is incidentally paramount in the context of this book as the more verifiable a task, the more we can improve it, and writing is extremely unverifiable. It is knowing this fact, and also knowing that a textbook is still a human-led endeavor, that I dictated Chapters 1 to 7 and edited every chapter,1 while Chapters 8, 10, and 11 and Appendix D started from drafts assembled by Claude from my notes and sources and then went through many rounds of my rewriting; by token count about half of the prose was first typed by a model, and every sentence has been edited by me. Beyond those drafts, the main contributions of Codex/Claude to this project were:
- helping me plan out the structure
- giving me the initial boilerplate/skeleton scaffold of the textbook itself
- creating the diagrams and equations, since this is much more efficient, and does not require the same human creativity as writing English (lower-entropy)
Target Audience
I wrote this book with the intent to cater to the largest audience possible. With that in mind, I increase difficulty as a function of the chapters such that if you are new to RLVR, you are best served in the beginning. If you are already experienced, you will gain the most from the later chapters.
How to Use This Book
Although the chapters do minimally build off of each other, they can still be read alone. Feel free to use the search function on the web version or Command F on the PDF to find what you wish directly. If you are new to RLVR, start with Chapter 1, Chapter 2, and Chapter 7; if you build training systems, Chapters 4, 5, and 10; and for frontier research, Chapters 9 to 11. The citations are plentiful to facilitate further research if there’s a specific theme which captivates you :).
Changelog
- 2026-04-16: Officially announced v0 of the book!
- 2026-04-19: Used my ML review textbook skill to refine each chapter and fix erratum.
- 2026-05-29: Expanded Chapter 9 with an end-to-end OLMo 3 Think training walkthrough.
- 2026-06-10: Reframed Chapter 10 around agentic harnesses and corrected the DeepSWE case study.
- 2026-07-18: Rebuilt Chapter 11 as an explicit RLVR research agenda.
- 2026-08-15: Added the ML textbook review skill used for the second review pass.
- 2026-09-04: Revised Chapters 1 and 2 and normalized citation placement across the book.
- 2026-09-05: Revised Chapter 3’s process reward explanations.
- 2026-09-22: Revised Chapter 4’s verifier explanations and Chapter 5’s reward shaping.
- 2026-09-23: Revised Chapter 6’s test time verification and Chapter 7’s reward hacking claims, and added Chapter 8 on optimization pressure in the wild.
- 2026-09-24: Revised Chapter 9’s frontier recipe, expanded Chapter 10 with multimodal verifiers, agent scaffolds, and environment synthesis, rebuilt Chapter 11 around verification, and added Appendix D with research ideas.
- 2026-09-25: Replaced ten Escher openers with higher-resolution scans and fixed PDF figure placement.
- 2026-09-26: Audited the whole book into ROADMAP.md, added the SemiAnalysis compute charts and redrawn figures to Chapters 7 and 11, replaced two Escher openers, added a license, and released v1.
Acknowledgments
I shamelessly take inspiration from Nathan Lambert’s RLHF book, and I am well aware that his textbook treats the subject of RLVR in detail; notwithstanding, as he notes himself, this particular sub-field of ML is evolving so fast that much of the RLHF book’s RLVR content will become outdated, and this book is intended to maintain pace with progress.
I also acknowledge the wonderful developers of Excalidraw, which I used for this book’s figures. Thanks to M.C. Escher for being the artistic soul of the book.2 Thanks to Simon Boehm for creating amazing educational content and establishing the target I strive to reach (same for Colah from distillpub)! Lastly, thanks to the quarto devs for making the software this book uses!
GitHub Contributors
Citation
You can cite this book directly with this BibTeX.
@online{kyars2026rlvrbook,
title = {Reinforcement Learning from Verifiable Rewards},
author = {Kyars, Kian},
year = {2026},
url = {https://rlvrbook.com},
}