Search Agentik

CtrlK

Glossary · Foundations

Mechanistic interpretability

Mechanistic interpretability is a research approach that tries to reverse engineer neural networks into understandable components, identifying the specific neurons, features and circuits that implement a model's behaviors.

Mechanistic interpretability sits in the Foundations part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.