Search Agentik

CtrlK

Glossary · Evaluation & safety

LLM-as-a-judge

LLM-as-a-judge is an evaluation method in which a language model scores or compares the outputs of another model against a rubric, offering scalable grading that should be checked against human judgments for reliability.

LLM-as-a-judge sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.

Also called LLM judge, Model-graded evaluation.