This project is suspended: no new articles or editions will be published. The archive stays available.

RLHF

RLHF (reinforcement learning from human feedback) is the training step that turns a raw language model into a helpful chatbot, using human rankings of its answers to shape its behavior. This hub explains how RLHF works, why it matters for AI safety and usefulness, and the news around it. Expect clear, jargon-free explainers.

Combine with: