LLM-as-a-Judge
Also called model-graded evaluation, LLM evaluator, autorater.
LLM-as-a-judge is the use of a strong language model to grade, score, or compare the outputs of models against criteria, as a scalable substitute for human evaluation.
Description
Zheng et al. found strong model judges agreed with human preferences at rates comparable to agreement between humans, while documenting biases toward the first-listed answer, longer answers, and the judge's own outputs.
Sources
- Zheng et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.
Cite this entry
Protologue. (2026). LLM-as-a-Judge. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0050). https://protologue.com/t/llm-as-a-judge/
BibTeX
@misc{protologue_llm_as_a_judge,
title = {LLM-as-a-Judge},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0050},
url = {https://protologue.com/t/llm-as-a-judge/}
}