protologue

Direct Preference Optimization

Also called DPO.

Direct preference optimization (DPO) aligns a language model to human preferences by training directly on preferred-versus-rejected response pairs, without fitting a separate reward model or running reinforcement learning.

Description

DPO reframes the RLHF objective as a simple classification-style loss over preference pairs. It became a common alternative to RLHF because it is stable and cheap to run.

Sources

  1. Rafailov et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Cite this entry

Protologue. (2026). Direct Preference Optimization. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0012). https://protologue.com/t/direct-preference-optimization/

BibTeX
@misc{protologue_direct_preference_optimization,
  title = {Direct Preference Optimization},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0012},
  url = {https://protologue.com/t/direct-preference-optimization/}
}

Markdown JSON