Glossary · Evaluation & safety
LLM-as-a-judge
LLM-as-a-judge is an evaluation method in which a language model scores or compares the outputs of another model against a rubric, offering scalable grading that should be checked against human judgments for reliability.
LLM-as-a-judge sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called LLM judge, Model-graded evaluation.