Back to glossary

Evaluation dataset

Değerlendirme veri kümesi

D4

An evaluation dataset is a defined collection of inputs or reference information used to assess a system for a stated purpose.

Review status: 2026-11-26

Technical explanation

Its documented accuracy, representativeness, relevance, and suitability help define what an evaluation result can describe.

Conceptual boundaries

An evaluation dataset is not automatically representative of deployment conditions; when training, validation, and testing roles are specified, those roles should not be silently conflated.

Provider-neutral example

A service team can keep a labelled set of past, permissioned support requests aside to examine whether a routing model meets its stated routing criteria.

Limitations

Documentation or use as known reference data does not by itself establish that a dataset is representative or that every label is uncontested.

Related concepts

Atomic claims and evidence

  1. 1.1NIST AI 600-1, Generative Artificial Intelligence Profile — D4 evidence slice
    Source
    NIST AI 600-1, Generative Artificial Intelligence Profile — D4 evidence slice
    Source role
    Authoritative source
    Exact locator
    MAP 2.3, MP-2.3-002, printed p. 28
    Supported claim
    NIST recommends reviewing and documenting the accuracy, representativeness, relevance, and suitability of data used at different AI lifecycle stages.
    Last verification
    Review due
    Scope limitation
    This guidance does not certify any dataset as representative merely because it is documented.
  2. 2.1NIST AI 600-1, Generative Artificial Intelligence Profile — D4 evidence slice
    Source
    NIST AI 600-1, Generative Artificial Intelligence Profile — D4 evidence slice
    Source role
    Authoritative source
    Exact locator
    MAP 2.3 action text, printed p. 27
    Supported claim
    NIST identifies known ground-truth data as one comparison input among varied evaluation methods.
    Last verification
    Review due
    Scope limitation
    Known reference data can support a task-specific comparison; it does not establish that all labels are complete or uncontested.
  3. 3.1Regulation (EU) 2024/1689 (Artificial Intelligence Act)
    Source
    Regulation (EU) 2024/1689 (Artificial Intelligence Act)
    Source role
    Authoritative source
    Exact locator
    Article 3(29)-(32)
    Supported claim
    The EU AI Act distinguishes training data used to fit learnable parameters, validation data used to evaluate a trained system and tune its non-learnable parameters and learning process, and testing data used to independently evaluate the system before it is placed on the market or put into service.
    Last verification
    Review due
    Scope limitation
    These are the Act's defined roles and do not require every technical workflow to use identical dataset splits.