Multimodality is a model's ability to handle several content types (text, image, audio, video) within one exchange: describing a photo, analysing a scanned document, commenting on a chart. Recent foundation models are natively multimodal, opening uses impossible with text alone.
In practice at Gensai
Gensai combines modalities per project: voice and text for voice agents, image and text for document analysis.