Glossary · Evaluation & safety
Interpretability
Interpretability is the degree to which humans can understand the internal workings of an AI model, such as which features or components drive its behavior, rather than only observing its inputs and outputs.
Interpretability sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.