Glossary · Foundations
Multi-head attention
Multi-head attention runs several attention operations in parallel, each learning to focus on different relationships in the data, then combines their outputs. It is a core component of transformers.
Multi-head attention sits in the Foundations part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.