Glossary · Models
Vision-language model
A vision-language model is a multimodal model that takes both images and text as input, enabling tasks such as describing pictures, answering questions about images, and reading documents.
Vision-language model sits in the Models part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called VLM.