Back to glossary

Multimodal model

Çok modlu model

D2

A multimodal model processes and relates information from multiple sensory modalities, such as vision and touch.

Review status: 2026-11-26

Technical explanation

It processes and relates information from more than one sensory modality, such as vision and touch.

Conceptual boundaries

Multimodal describes the modalities a model processes; it does not by itself guarantee robustness.

Provider-neutral example

A workflow can provide a photo and a text question to one model and ask for a draft description that a reviewer checks.

Limitations

Information from multiple modalities does not necessarily make a model robust to an adversarial change in one modality.

Related concepts

Atomic claims and evidence

  1. 1.1NIST AI 100-2e2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations
    Source
    NIST AI 100-2e2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations
    Source role
    Authoritative source
    Exact locator
    Appendix A p. 110: multimodal models
    Supported claim
    A multimodal model processes and relates information from multiple sensory modalities that represent primary human communication and sensation channels.
    Last verification
    Review due
    Scope limitation
    The cited definition uses sensory modalities and examples such as vision and touch; it does not prescribe a fixed list of modalities.
  2. 2.1NIST AI 100-2e2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations
    Source
    NIST AI 100-2e2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations
    Source role
    Authoritative source
    Exact locator
    Section 4.2.3 pp. 58-59: Multimodal Models
    Supported claim
    Redundancy across modalities does not necessarily make a multimodal model robust to adversarial perturbations of a single modality.
    Last verification
    Review due
    Scope limitation
    This is a risk finding about adversarial perturbations, not a claim that every multimodal model is vulnerable in the same way.