In plain English
Multimodal AI describes systems that can process and generate more than one type of data, such as text, images, audio, and video, within a single model. This lets a model take a photo and a question together, or turn a spoken request into a written answer. It moves AI closer to the way people naturally mix senses and media.
A worked example
You photograph a fridge full of ingredients, ask what you can cook, and a multimodal model reads the image and replies with recipes.
Common confusion
Multimodal does not mean separate tools bolted together. A truly multimodal model handles the different data types within one system rather than passing files between specialised models.

