Back to glossary

Benchmark

Kıyaslama testi

D4

A benchmark is a defined comparative evaluation arrangement that combines tasks or data, a protocol, and one or more metrics.

Review status: 2027-08-28

Technical explanation

It lets people compare results under stated tasks and scoring conditions.

Conceptual boundaries

A benchmark is not a single metric; its results describe the stated arrangement rather than every operational setting.

Provider-neutral example

Several teams can run the same versioned set of classification tasks with the same scoring rules and compare their reported results alongside the task limitations.

Limitations

Performance can vary materially across individual tasks, and laboratory measurements can differ from risks in operational settings.

Related concepts

Atomic claims and evidence

  1. 1.1Hendrycks et al., Measuring Massive Multitask Language Understanding
    Source
    Hendrycks et al., Measuring Massive Multitask Language Understanding
    Source role
    Authoritative source
    Exact locator
    Abstract: 57-task test, task-level results, and shortcomings
    Supported claim
    The MMLU study presents a 57-task test for comparing text-model multitask accuracy and reports task-level variation in model performance.
    Last verification
    Review due
    Scope limitation
    This describes one benchmark and its results; it does not establish a universal benchmark design.
  2. 2.1Liang et al., Holistic Evaluation of Language Models
    Source
    Liang et al., Holistic Evaluation of Language Models
    Source role
    Authoritative source
    Exact locator
    Abstract; Sections 2.3-2.4: scenarios, metrics, and adaptations
    Supported claim
    HELM describes evaluation as a space of scenarios, metrics, and adaptations rather than a single score.
    Last verification
    Review due
    Scope limitation
    This is the research framework's taxonomy and does not require every benchmark to cover all possible scenarios.
  3. 3.1NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source
    NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source role
    Authoritative source
    Exact locator
    Section 1.2.1, Risk in real-world settings, printed p. 6
    Supported claim
    NIST notes that laboratory measurements can differ from risks emerging in operational settings.
    Last verification
    Review due
    Scope limitation
    This does not make benchmark evidence useless; it limits what it can establish about later operation.