Glossary · Evaluation & safety
Benchmark
A benchmark is a standardized dataset and scoring method used to compare AI models on the same task, such as reasoning, coding or knowledge questions, so results can be reported consistently across models and versions.
Benchmark sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called AI benchmark.