Glossary · Infrastructure
Inference server
An inference server is software that loads AI models into memory and answers prediction or generation requests over a network, with features such as batching, multi-model hosting and metrics.
Inference server sits in the Infrastructure part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called Model server.