Multimodal
A model that can handle more than just text: images, audio, or video as input or output.
It lets AI read a screenshot, describe a photo, or listen to a recording, not just chat.
Frequently asked questions
What does multimodal mean?
Multimodal describes a model that can handle more than just text, such as images, audio, or video, as input or output. It can read a screenshot, describe a photo, or listen to a recording, not only chat in words.
Can I show an AI a picture or a document instead of typing everything?
With a multimodal model, yes: you can upload an image, a screenshot, or a photo and ask questions about it directly. That is exactly what multimodal means, handling more than plain text.
What are practical uses of multimodal AI for a small business?
You might have it read a photographed receipt, describe a product image, pull text out of a screenshot, or summarize a voice memo. It is useful anywhere the information lives in a picture or a recording rather than typed text.
New to all this? Start with what an AI agent really is, browse the full glossary, or explore the learning hub.
← Back to the glossary