Search Agentik

CtrlK

Glossary · Models

Vision-language model

A vision-language model is a multimodal model that takes both images and text as input, enabling tasks such as describing pictures, answering questions about images, and reading documents.

Vision-language model sits in the Models part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.

Also called VLM.