Metric
Metrik
D4
A metric operationalizes a stated criterion for assessing a system or its output.
Review status: 2027-08-28
Technical explanation
A metric operationalizes a stated criterion; its choice and interpretation can omit context, be gamed, or hide group differences.
Conceptual boundaries
A metric is one component of an evaluation arrangement that can also specify scenarios and adaptations.
Provider-neutral example
For a routing task, a team can calculate the proportion of requests sent to the correct queue and inspect which request types the calculation hides.
Limitations
Metrics can be oversimplified, gamed, or poorly aligned with affected groups and contexts.
Related concepts
Atomic claims and evidence
1.1NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
- Source
- NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
- Source role
- Authoritative source
- Exact locator
- Section 1.2.1, "Availability of reliable metrics," printed p. 6
- Supported claim
- NIST identifies lack of consensus on robust and verifiable measurement methods and differences across use cases as an AI risk-measurement challenge.
- Last verification
- Review due
- Scope limitation
- This does not mean metrics are unusable; it means their adequacy must be considered in context.
2.1NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
- Source
- NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
- Source role
- Authoritative source
- Exact locator
- Section 1.2.1, "Availability of reliable metrics," printed p. 6
- Supported claim
- NIST warns that measurement approaches can be oversimplified, gamed, relied on unexpectedly, or fail to account for differences in groups and contexts.
- Last verification
- Review due
- Scope limitation
- This is a warning about interpretation and design, not a claim that every metric is misleading.
3.1Liang et al., Holistic Evaluation of Language Models
- Source
- Liang et al., Holistic Evaluation of Language Models
- Source role
- Authoritative source
- Exact locator
- Sections 2.3-2.4: desiderata, scenarios, adaptations, and metrics
- Supported claim
- HELM treats metrics as operationalizations of desired properties and structures evaluation runs with scenarios, adaptations, and metrics.
- Last verification
- Review due
- Scope limitation
- This describes HELM's evaluation framework and does not prescribe one metric set for every system.