Back to glossary

AI evaluation

YZ değerlendirmesi

D4

AI evaluation is a structured process that produces evidence about how an AI system performs against stated criteria in a stated context.

Review status: 2027-08-28

Technical explanation

It documents qualitative or quantitative performance or assurance criteria under conditions similar to the stated deployment setting.

Conceptual boundaries

Evidence from one setting does not establish how a system will perform in every operational setting.

Provider-neutral example

A team can test a document-classification system on defined examples, record error types and thresholds, and decide whether its stated internal workflow is ready for a pilot.

Limitations

Laboratory measurements can differ from risks that emerge in operational settings.

Related concepts

Atomic claims and evidence

  1. 1.1NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source
    NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source role
    Authoritative source
    Exact locator
    Table 3, MEASURE 2.3, printed p. 29
    Supported claim
    NIST's MEASURE function calls for qualitative or quantitative performance or assurance criteria to be demonstrated for conditions similar to deployment settings and documented.
    Last verification
    Review due
    Scope limitation
    This is risk-management guidance, not a requirement that every evaluation use the same metric or protocol.
  2. 2.1NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source
    NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source role
    Authoritative source
    Exact locator
    Section 1.2.1, Risk in real-world settings, printed p. 6
    Supported claim
    NIST notes that laboratory measurements can differ from risks emerging in operational settings.
    Last verification
    Review due
    Scope limitation
    This does not make pre-deployment evaluation useless; it limits what it can establish about later use.