← AI Glossary

Multimodality

Language, voice & vision

Multimodality is a model's ability to handle several content types (text, image, audio, video) within one exchange: describing a photo, analysing a scanned document, commenting on a chart. Recent foundation models are natively multimodal, opening uses impossible with text alone.

In practice at Gensai

Gensai combines modalities per project: voice and text for voice agents, image and text for document analysis.