新书发布:RLHF 后训练教科书,涵盖 PPO 与系统设计
原标题:5 useful things you'll learn in my new post-training textbook (shipping now!)
AI 摘要
Interconnects AI 的作者 Nathan Lambert 发布了新书《Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs》,由 Manning 出版,旨在系统讲解 LLM 后训练(post-training)的关键方法、直觉与历史。书中涵盖 RL 算法(如 PPO、GSPO)、系统设计、以及 DAPO 等变体,并附带 12 小时课程和代码库。该书在 Manning 上五折促销至 8 月 19 日,并已开始发货。
正文节选
5 useful things you'll learn in my new post-training textbook (shipping now!) Reinforcement Learning from Human Feedback is coming to a neolab near you. Housekeeping: No voiceover on another quick “launch” post. More essays soon! After a few long years of finding time to document my lessons from training open models, my post-training book is done! It’s published by Manning, under the title Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs. Telling the story of the book