Glossary · Evaluation & safety
Evaluation
An evaluation is a structured test that measures how well an AI model or system performs on a defined task, using prepared inputs, expected outputs or scoring rules, and metrics that can be tracked over time.
Evaluation sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called Eval, LLM evaluation, Model evaluation.