◈ AI GLOSSARY ◈

Multimodal

A model that can handle more than just text: images, audio, or video as input or output.

WHY IT MATTERS

It lets AI read a screenshot, describe a photo, or listen to a recording, not just chat.

RELATED TERMS

Frequently asked questions

What does multimodal mean?

Multimodal describes a model that can handle more than just text, such as images, audio, or video, as input or output. It can read a screenshot, describe a photo, or listen to a recording, not only chat in words.

Can I show an AI a picture or a document instead of typing everything?

With a multimodal model, yes: you can upload an image, a screenshot, or a photo and ask questions about it directly. That is exactly what multimodal means, handling more than plain text.

What are practical uses of multimodal AI for a small business?

You might have it read a photographed receipt, describe a product image, pull text out of a screenshot, or summarize a voice memo. It is useful anywhere the information lives in a picture or a recording rather than typed text.

New to all this? Start with what an AI agent really is, browse the full glossary, or explore the learning hub.

← Back to the glossary