Benchmark
Kıyaslama testi
D4
A benchmark is a defined comparative evaluation arrangement that combines tasks or data, a protocol, and one or more metrics.
Review status: 2027-08-28
Technical explanation
It lets people compare results under stated tasks and scoring conditions.
Conceptual boundaries
A benchmark is not a single metric; its results describe the stated arrangement rather than every operational setting.
Provider-neutral example
Several teams can run the same versioned set of classification tasks with the same scoring rules and compare their reported results alongside the task limitations.
Limitations
Performance can vary materially across individual tasks, and laboratory measurements can differ from risks in operational settings.
Related concepts
Atomic claims and evidence
1.1Hendrycks et al., Measuring Massive Multitask Language Understanding
- Source
- Hendrycks et al., Measuring Massive Multitask Language Understanding
- Source role
- Authoritative source
- Exact locator
- Abstract: 57-task test, task-level results, and shortcomings
- Supported claim
- The MMLU study presents a 57-task test for comparing text-model multitask accuracy and reports task-level variation in model performance.
- Last verification
- Review due
- Scope limitation
- This describes one benchmark and its results; it does not establish a universal benchmark design.
2.1Liang et al., Holistic Evaluation of Language Models
- Source
- Liang et al., Holistic Evaluation of Language Models
- Source role
- Authoritative source
- Exact locator
- Abstract; Sections 2.3-2.4: scenarios, metrics, and adaptations
- Supported claim
- HELM describes evaluation as a space of scenarios, metrics, and adaptations rather than a single score.
- Last verification
- Review due
- Scope limitation
- This is the research framework's taxonomy and does not require every benchmark to cover all possible scenarios.
3.1NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
- Source
- NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
- Source role
- Authoritative source
- Exact locator
- Section 1.2.1, Risk in real-world settings, printed p. 6
- Supported claim
- NIST notes that laboratory measurements can differ from risks emerging in operational settings.
- Last verification
- Review due
- Scope limitation
- This does not make benchmark evidence useless; it limits what it can establish about later operation.