Glossary · Evaluation & safety
Offline evaluation
Offline evaluation is the testing of an AI system on fixed, prepared datasets before or outside of production, allowing controlled comparisons between models, prompts or versions without affecting real users.
Offline evaluation sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.