AI Glossary
Multimodal
What is Multimodal?
AI systems that can process and generate more than one type of data — for example, text, images, audio, and video within a single model. GPT-4o and Gemini are multimodal models. Multimodal capability enables use cases such as analysing a chart uploaded as an image, or transcribing and summarising a meeting recording.
Example in practice
A product manager uploading a screenshot of a competitor's pricing page and asking an AI to extract and compare the pricing tiers is using multimodal capability — the model processes both the image and the follow-up text prompt together.
Learn more
See Multimodal applied in a professional context through this free course.
AI Fundamentals for Professionals →