AI Glossary
Vision Language Model (VLM)
What is Vision Language Model (VLM)?
A multimodal AI model that can process both images and text, enabling tasks like image captioning, visual question answering, and document analysis. GPT-4o and Claude 3 are examples of vision language models. VLMs are increasingly used in business applications that involve processing invoices, diagrams, photographs, and mixed-format documents.
Example in practice
An accounts payable team that uploads scanned invoices directly to an AI assistant for automatic data extraction — without any OCR pre-processing step — is using a VLM's ability to read and understand document images natively.
Learn more
See Vision Language Model (VLM) applied in a professional context through this free course.
AI Fundamentals for Professionals →