# Reinforcement Learning from Human Feedback

> Reinforcement learning from human feedback (RLHF) trains a language model to produce outputs people prefer, by learning a reward model from human comparisons and optimizing the model against it.

- Identifier: PTL-0011
- Category: Foundations
- Canonical URL: https://protologue.com/t/rlhf/
- Also known as: RLHF, preference tuning
- Introduced: 2022

## Description

InstructGPT showed that RLHF made a much smaller model preferred over a far larger base model. RLHF shapes how models respond to prompts, including their helpfulness and refusals, and is linked to failure modes such as sycophancy. Direct preference optimization is a widely used simpler alternative.

## Narrower terms

- [Direct Preference Optimization](https://protologue.com/t/direct-preference-optimization/)

## Related terms

- [Instruction Tuning](https://protologue.com/t/instruction-tuning/)
- [Constitutional AI](https://protologue.com/t/constitutional-ai/)
- [Sycophancy](https://protologue.com/t/sycophancy/)

## Sources

- Ouyang et al. (2022). Training language models to follow instructions with human feedback. https://arxiv.org/abs/2203.02155

## Cite this entry

Protologue. (2026). Reinforcement Learning from Human Feedback. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0011). https://protologue.com/t/rlhf/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
