protologue

LLM-as-a-Judge

Also called model-graded evaluation, LLM evaluator, autorater.

LLM-as-a-judge is the use of a strong language model to grade, score, or compare the outputs of models against criteria, as a scalable substitute for human evaluation.

Description

Zheng et al. found strong model judges agreed with human preferences at rates comparable to agreement between humans, while documenting biases toward the first-listed answer, longer answers, and the judge's own outputs.

Sources

  1. Zheng et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Cite this entry

Protologue. (2026). LLM-as-a-Judge. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0050). https://protologue.com/t/llm-as-a-judge/

BibTeX
@misc{protologue_llm_as_a_judge,
  title = {LLM-as-a-Judge},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0050},
  url = {https://protologue.com/t/llm-as-a-judge/}
}

Markdown JSON