Glossary · Foundations
Mechanistic interpretability
Mechanistic interpretability is a research approach that tries to reverse engineer neural networks into understandable components, identifying the specific neurons, features and circuits that implement a model's behaviors.
Mechanistic interpretability sits in the Foundations part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.