Back to glossary

Human evaluation

İnsan değerlendirmesi

D4

Human evaluation is structured human judgment of a system or its outputs under a stated task and protocol.

Review status: 2026-11-26

Technical explanation

Raters can use annotation guidelines, rating scales, and multiple annotators so that judgments and their variation can be recorded.

Conceptual boundaries

Human evaluation is not unrecorded opinion and is not automatically free of disagreement.

Provider-neutral example

Using an agreed rubric, two reviewers can independently judge whether a generated summary preserves required facts, then record agreement and unresolved cases.

Limitations

Human annotation can contain ambiguity and disagreement, so a protocol should record rather than conceal variation.

Related concepts

Atomic claims and evidence

  1. 1.1NIST AI 600-1, Generative Artificial Intelligence Profile — D4 evidence slice
    Source
    NIST AI 600-1, Generative Artificial Intelligence Profile — D4 evidence slice
    Source role
    Authoritative source
    Exact locator
    MAP 2.3 action text, printed p. 27
    Supported claim
    NIST lists human oversight and automated evaluation among varied methods that can be used when evaluating information integrity.
    Last verification
    Review due
    Scope limitation
    The source does not equate human oversight with a single rating method or guarantee agreement among people.
  2. 2.1NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source
    NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
    Source role
    Authoritative source
    Exact locator
    Section 3, printed p. 12
    Supported claim
    NIST states that human judgment should be used in choosing metrics and their threshold values for trustworthiness characteristics.
    Last verification
    Review due
    Scope limitation
    This supports explicit human judgment in the process; it does not establish that a human judgment is objective.
  3. 3.1Liang et al., Holistic Evaluation of Language Models
    Source
    Liang et al., Holistic Evaluation of Language Models
    Source role
    Authoritative source
    Exact locator
    Section 8.5.1: annotation guidelines, rating scales, and multiple annotators
    Supported claim
    HELM describes human evaluation practices using annotation guidelines, rating scales, and multiple annotators.
    Last verification
    Review due
    Scope limitation
    This describes reported evaluation practices and does not prescribe one protocol for every task.
  4. 4.1Aroyo and Welty, Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation
    Source
    Aroyo and Welty, Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation
    Source role
    Authoritative source
    Exact locator
    Abstract; "The Seven Myths" and "Disagreement Is Bad," printed pp. 16-17
    Supported claim
    Aroyo and Welty describe ambiguity and disagreement as information-bearing features of human annotation.
    Last verification
    Review due
    Scope limitation
    This does not say that every disagreement is desirable or that human evaluation has no usable conclusion.