Human evaluation
İnsan değerlendirmesi
D4
Human evaluation is structured human judgment of a system or its outputs under a stated task and protocol.
Review status: 2026-11-26
Technical explanation
Raters can use annotation guidelines, rating scales, and multiple annotators so that judgments and their variation can be recorded.
Conceptual boundaries
Human evaluation is not unrecorded opinion and is not automatically free of disagreement.
Provider-neutral example
Using an agreed rubric, two reviewers can independently judge whether a generated summary preserves required facts, then record agreement and unresolved cases.
Limitations
Human annotation can contain ambiguity and disagreement, so a protocol should record rather than conceal variation.
Related concepts
Atomic claims and evidence
1.1NIST AI 600-1, Generative Artificial Intelligence Profile — D4 evidence slice
- Source
- NIST AI 600-1, Generative Artificial Intelligence Profile — D4 evidence slice
- Source role
- Authoritative source
- Exact locator
- MAP 2.3 action text, printed p. 27
- Supported claim
- NIST lists human oversight and automated evaluation among varied methods that can be used when evaluating information integrity.
- Last verification
- Review due
- Scope limitation
- The source does not equate human oversight with a single rating method or guarantee agreement among people.
2.1NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
- Source
- NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) — D4 evidence slice
- Source role
- Authoritative source
- Exact locator
- Section 3, printed p. 12
- Supported claim
- NIST states that human judgment should be used in choosing metrics and their threshold values for trustworthiness characteristics.
- Last verification
- Review due
- Scope limitation
- This supports explicit human judgment in the process; it does not establish that a human judgment is objective.
3.1Liang et al., Holistic Evaluation of Language Models
- Source
- Liang et al., Holistic Evaluation of Language Models
- Source role
- Authoritative source
- Exact locator
- Section 8.5.1: annotation guidelines, rating scales, and multiple annotators
- Supported claim
- HELM describes human evaluation practices using annotation guidelines, rating scales, and multiple annotators.
- Last verification
- Review due
- Scope limitation
- This describes reported evaluation practices and does not prescribe one protocol for every task.
4.1Aroyo and Welty, Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation
- Source
- Aroyo and Welty, Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation
- Source role
- Authoritative source
- Exact locator
- Abstract; "The Seven Myths" and "Disagreement Is Bad," printed pp. 16-17
- Supported claim
- Aroyo and Welty describe ambiguity and disagreement as information-bearing features of human annotation.
- Last verification
- Review due
- Scope limitation
- This does not say that every disagreement is desirable or that human evaluation has no usable conclusion.