RLHF
Reinforcement Learning from Human Feedback: humans rank outputs, a reward model learns those preferences, and the LLM is optimized against it. The technique that made raw predictors into helpful assistants.
Reinforcement Learning from Human Feedback: humans rank outputs, a reward model learns those preferences, and the LLM is optimized against it. The technique that made raw predictors into helpful assistants.