Direct Preference Optimization
Also called DPO.
Direct preference optimization (DPO) aligns a language model to human preferences by training directly on preferred-versus-rejected response pairs, without fitting a separate reward model or running reinforcement learning.
Description
DPO reframes the RLHF objective as a simple classification-style loss over preference pairs. It became a common alternative to RLHF because it is stable and cheap to run.
Sources
- Rafailov et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model.
Cite this entry
Protologue. (2026). Direct Preference Optimization. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0012). https://protologue.com/t/direct-preference-optimization/
BibTeX
@misc{protologue_direct_preference_optimization,
title = {Direct Preference Optimization},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0012},
url = {https://protologue.com/t/direct-preference-optimization/}
}