跳转至

文章背景与核心概要

大语言模型(LLM)目前在模型评估中扮演着多重角色——既是接受评分的被试者、其他模型的裁判,又是人类内容的评估者。然而,标准的评估实践往往掩盖了各个组件的独立贡献。本文将拉斯施测量理论(Rasch Measurement Theory, RMT)引入大语言模型评估中,作为一种严谨的方法,将定序评分分解为统一量表上的可分离维度并诊断偏见。

通过以“测量仇恨言论(Measuring Hate Speech)”语料库作为案例研究,作者评估了9个大语言模型,发现它们与人类评估者在严重程度、题目校准、问题顺序鲁棒性、目标身份敏感度以及量表使用习惯上存在系统性差异。作者主张在被试者、裁判和评估者等多种范式中全面整合RMT方法,以提升大模型评估的科学性与透明度。


Rating the Raters: Rasch Measurement Theory for LLM Evaluation

Rating the Raters: Rasch Measurement Theory for LLM Evaluation

License: CC BY 4.0

License: CC BY 4.0

Summary

Summary

大语言模型(LLM)目前在评估中担任多重角色——作为被试者、其他模型的裁判以及人类内容的评估者。然而,标准的评估实践往往掩盖了每个组件的个体贡献。本文将拉斯施测量理论(Rasch Measurement Theory, RMT)引入LLM评估中,作为一种严谨的方法,能够将定序评分分解为统一量表上的可分离维度并诊断偏见。以“测量仇恨言论(Measuring Hate Speech)”语料库为案例研究,作者评估了9个LLM,并发现它们与人类评估者在严重程度、题目校准、问题顺序鲁棒性、目标身份敏感度和量表使用方面存在系统性差异。作者主张将RMT整合到被试者、裁判和评估者等各种评估范式中。

Large Language Models (LLMs) currently serve multiple roles in evaluation—as examinees, judges of other models, and raters of human content. However, standard evaluation practices often obscure the individual contributions of each component. This paper introduces Rasch Measurement Theory (RMT) to LLM evaluation as a rigorous method to decompose ordinal ratings into separable facets on a common scale and diagnose biases. Using the Measuring Hate Speech corpus as a case study, the authors evaluate nine LLMs and discover systematic differences from human raters in severity, item calibration, question-order robustness, target-identity sensitivity, and scale usage. The authors advocate for integrating RMT across examinee, judge, and rater paradigms.


Metadata


Metadata

  • arXiv ID: arXiv:2608.27463
  • Subject: 人工智能 (cs.AI)
  • Authors: Pratik S. Sachdeva, Nathan Boudol
  • Status: 已被 EMNLP 2026 录用
  • Submitted: 2026年7月14日 (v1);最后修订:2026年9月1日 (v2)
  • DOI: 10.48550/arXiv.2608.27463
  • arXiv ID: arXiv:2608.27463
  • Subject: Artificial Intelligence (cs.AI)
  • Authors: Pratik S. Sachdeva, Nathan Boudol
  • Status: Accepted to EMNLP 2026
  • Submitted: 14 Jul 2026 (v1); Last revised: 1 Sep 2026 (v2)
  • DOI: 10.48550/arXiv.2608.27463

Abstract


Abstract

大语言模型如今在评估中处于各个环节:作为在基准测试中被评分的被试者、其他模型输出的裁判,以及人类生成内容的评估者。每种范式都可以被视为一个测量问题,即通过裁判或评估者,利用测量工具(例如基准测试)中的题目来探测对象的潜在属性。

LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by judges or raters.

标准的评估实践往往忽略了各个核心组件对最终结果的贡献,从而限制了我们对被测量内容的理解。拉斯施测量理论(RMT)非常适合解决这一问题。RMT能够将定序评分分解为统一量表上的可分离维度。此外,它还提供了一套诊断工具,可以识别校准不当的测量结果和评估者偏见。

Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured. Rasch measurement theory (RMT) is well-suited to this problem. RMT decomposes ordinal ratings into separable facets on a common scale. It further provides a battery of diagnostics that can identify miscalibrated measurements and rater biases.

我们展示了将RMT应用于“LLM作为评估者”范式的案例研究,该研究所基于的语料库本身就是根据RMT构建的。我们将一系列多面拉斯施模型(many-facet Rasch models)拟合到来自9个跨越不同系列和能力水平的LLM的标注数据中。我们的分析表明,LLM在严重程度、题目层面的校准、问题顺序鲁棒性、目标身份敏感度以及评分量表的使用上,与人类评估者存在系统性差异,而标准的评估实践大体上会掩盖这些差异。总体而言,我们认为RMT应当纳入评估LLM作为被试者、裁判和评估者范式的工具箱中。

We present a case study of RMT applied to the LLM-as-rater paradigm using the Measuring Hate Speech corpus, whose construct was itself built under RMT. We fit a series of many-facet Rasch models to annotations from nine LLMs spanning families and capability levels. Our analyses show that LLMs systematically differ from human raters in severity, item-level calibration, question-order robustness, target-identity sensitivity, and rating scale use, all of which standard evaluation practice would largely obscure. Overall, we argue that RMT belongs in the toolkit for evaluating LLM-as-examinee, -judge, and -rater paradigms.


Full-Text & Resources


Full-Text & Resources


References & External Tools


References & External Tools