An AI model trained to handle multiple modalities of data, such as text, images, and audio. This allows the model to understand and leverage information from various sources for better decision-making.