Signal & Noise
Menu

DPO (Direct Preference Optimization)

A preference-tuning method that trains directly on chosen-vs-rejected response pairs, replacing the more complex RLHF-with-PPO pipeline for most alignment work.

Related terms

← Back to the full glossary