Back to glossary

Metric

Metrik

D4

A metric operationalizes a stated criterion for assessing a system or its output.

Review status: 2027-08-28

Technical explanation

A metric operationalizes a stated criterion; its choice and interpretation can omit context, be gamed, or hide group differences.

Conceptual boundaries

A metric is one component of an evaluation arrangement that can also specify scenarios and adaptations.

Provider-neutral example

For a routing task, a team can calculate the proportion of requests sent to the correct queue and inspect which request types the calculation hides.

Limitations

Metrics can be oversimplified, gamed, or poorly aligned with affected groups and contexts.

Related concepts

Atomic claims and evidence

  1. 1.1NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source
    NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source role
    Authoritative source
    Exact locator
    Section 1.2.1, "Availability of reliable metrics," printed p. 6
    Supported claim
    NIST identifies lack of consensus on robust and verifiable measurement methods and differences across use cases as an AI risk-measurement challenge.
    Last verification
    Review due
    Scope limitation
    This does not mean metrics are unusable; it means their adequacy must be considered in context.
  2. 2.1NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source
    NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source role
    Authoritative source
    Exact locator
    Section 1.2.1, "Availability of reliable metrics," printed p. 6
    Supported claim
    NIST warns that measurement approaches can be oversimplified, gamed, relied on unexpectedly, or fail to account for differences in groups and contexts.
    Last verification
    Review due
    Scope limitation
    This is a warning about interpretation and design, not a claim that every metric is misleading.
  3. 3.1Liang et al., Holistic Evaluation of Language Models
    Source
    Liang et al., Holistic Evaluation of Language Models
    Source role
    Authoritative source
    Exact locator
    Sections 2.3-2.4: desiderata, scenarios, adaptations, and metrics
    Supported claim
    HELM treats metrics as operationalizations of desired properties and structures evaluation runs with scenarios, adaptations, and metrics.
    Last verification
    Review due
    Scope limitation
    This describes HELM's evaluation framework and does not prescribe one metric set for every system.