Search Agentik

CtrlK

Glossary · Evaluation & safety

Evaluation

An evaluation is a structured test that measures how well an AI model or system performs on a defined task, using prepared inputs, expected outputs or scoring rules, and metrics that can be tracked over time.

Evaluation sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.

Also called Eval, LLM evaluation, Model evaluation.