DPO (Direct Preference Optimization)
A preference-tuning method that trains directly on chosen-vs-rejected response pairs, replacing the more complex RLHF-with-PPO pipeline for most alignment work.
A preference-tuning method that trains directly on chosen-vs-rejected response pairs, replacing the more complex RLHF-with-PPO pipeline for most alignment work.