Glossary · Infrastructure
Model parallelism
Model parallelism is splitting a single model across multiple GPUs or machines, either by dividing its layers or its individual tensor operations, so models too large for one device can be trained or served.
Model parallelism sits in the Infrastructure part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.