Skip to main content
Deliberate AcademyProfessional AI Education
AI Glossary

Multimodal

What is Multimodal?

AI systems that can process and generate more than one type of data — for example, text, images, audio, and video within a single model. GPT-4o and Gemini are multimodal models. Multimodal capability enables use cases such as analysing a chart uploaded as an image, or transcribing and summarising a meeting recording.

Example in practice

A product manager uploading a screenshot of a competitor's pricing page and asking an AI to extract and compare the pricing tiers is using multimodal capability — the model processes both the image and the follow-up text prompt together.

Learn more

See Multimodal applied in a professional context through this free course.

AI Fundamentals for Professionals