跳转至

文章背景与核心概要

本文探讨了自然语言生成(NLG)领域中标准人类评估协议的可靠性问题。作者指出,目前广泛采用的收集并平均化李克特量表(Likert-scale)评分的做法,由于其潜在的假设存在缺陷,往往无法捕捉到人类真实的偏好。

通过应用效用理论,该研究证明了传统评估方法可能导致悖论性的结果,甚至出现与真实偏好方向相反的情况。为了解决这些局限性,作者提出了一种名为“系统级概率评估”(System-level Probabilistic Assessment, SPA)的方案,这是一种更稳健的评估协议,能够在传统方法失效的情况下成功恢复预期的模型排名。


人类评估中的真实性鸿沟

作者: Kawin Ethayarajh, Dan Jurafsky
发表于: EMNLP 2022 (arXiv:2205.11930)

The Authenticity Gap in Human Evaluation

Authors: Kawin Ethayarajh, Dan Jurafsky
Published: EMNLP 2022 (arXiv:2205.11930)


摘要

本文调查了自然语言生成(NLG)中标准人类评估协议的可靠性。作者认为,收集和平均化李克特量表评分的常见做法,由于存在有缺陷的潜在假设,往往无法捕捉到真实的人类偏好。通过应用效用理论,该研究证明了这些方法可能导致悖论性的结果,例如逆转实际偏好的方向。为了解决这些局限性,作者引入了系统级概率评估(SPA),这是一种更稳健的评估协议,能够在传统方法失效时成功恢复预期的模型排名。

Summary

This paper investigates the reliability of standard human evaluation protocols in Natural Language Generation (NLG). The authors argue that the common practice of collecting and averaging Likert-scale ratings often fails to capture true human preferences due to flawed underlying assumptions. By applying utility theory, the study demonstrates that these methods can lead to paradoxical results, such as reversing the direction of actual preference. To address these limitations, the authors introduce System-level Probabilistic Assessment (SPA), a more robust evaluation protocol that successfully recovers expected model rankings where traditional methods fail.


关键发现

  • 标准协议的缺陷: 目前的“黄金标准”——收集并平均化李克特评分——依赖于关于标注者的隐性假设,而这些假设在现实场景中经常被违背。
  • 李克特量表问题: 在特定情况下,使用李克特量表可以被证明会逆转真实人类偏好的方向。
  • 真实性鸿沟: 传统协议在评估故事生成等开放式任务时表现吃力。例如,在比较 GPT-3 模型时,标准方法无法在 curiedavinci 等模型之间发现显著差异。
  • 提出的解决方案(SPA): 作者提出了“系统级概率评估”(SPA)。当应用于故事生成任务时,SPA 成功地以统计学显著性恢复了 GPT-3 模型按规模排列的预期顺序。

Key Findings

  • The Flaw in Standard Protocols: The current "gold standard"—collecting and averaging Likert ratings—relies on implicit assumptions about annotators that are frequently violated in real-world scenarios.
  • The Likert Scale Problem: The use of Likert scales can, in specific instances, provably reverse the direction of true human preference.
  • The Authenticity Gap: Traditional protocols struggle to evaluate open-ended tasks like story generation. For example, when comparing GPT-3 models, standard methods failed to find significant differences between models like curie and davinci.
  • A Proposed Solution (SPA): The authors propose System-level Probabilistic Assessment (SPA). When applied to story generation tasks, SPA successfully recovers the expected ordering of GPT-3 models by size with statistical significance.

论文元数据

详情 信息
主要学科 计算与语言 (cs.CL)
次要学科 人工智能 (cs.AI), 机器学习 (cs.LG)
DOI 10.48550/arXiv.2205.11930
许可协议 license icon view license

Paper Metadata

Detail Information
Primary Subject Computation and Language (cs.CL)
Secondary Subjects Artificial Intelligence (cs.AI), Machine Learning (cs.LG)
DOI 10.48550/arXiv.2205.11930
License license icon view license

获取研究内容

Access the Research