文章背景与核心概要
在大模型(LLM)的生态系统中,模型正扮演着越来越多的角色:它们既是接受基准测试的“考生”、评估其他模型输出的“裁判”,又是对人类生成内容进行打分的“评委”。然而,传统的评估范式往往忽视了模型本身、基准测试题目以及评分量表等各个独立组件如何对最终结果产生影响。为了解决这一局限性,本文首次将心理计量学中的拉施测量理论(Rasch Measurement Theory, RMT)引入大模型评估中,为解构有序评分、诊断评委偏差提供了强有力的数学工具。
研究人员以基于 RMT 构建的《测量仇恨言论》(Measuring Hate Speech)语料库为案例,评估了来自不同模型家族和能力层级的 9 个大模型的标注表现。研究发现,大模型在严重程度、题目级校准、题目顺序鲁棒性、目标群体敏感性以及评分量表使用习惯等方面,与人类评委存在系统性的系统偏差。这些关键的偏差在传统评估实践中极易被掩盖,因此作者强烈主张将 RMT 纳入大模型作为考生、裁判和评委时的标准评估工具箱中。
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
Authors: Pratik S. Sachdeva, Nathan Boudol
Primary Subject: Artificial Intelligence (cs.AI)
arXiv ID: arXiv:2608.27463 [cs.AI]
Submitted: 14 July 2026
License: Creative Commons Attribution 4.0 International
Summary
大语言模型(LLM)目前在评估流水线中扮演着多重角色:它们既是接受基准测试评分的考生、评估其他模型输出的裁判,也是对人类生成内容进行打分的评委。传统的评估范式通常对这些角色进行简单化处理,忽略了各个独立组件(模型、基准测试项目和评分量表)如何对最终结果产生影响。
Large Language Models (LLMs) currently occupy multiple roles in evaluation pipelines: they act as examinees scored on benchmarks, judges evaluating other models' outputs, and raters of human-generated content. Traditional evaluation paradigms often treat these roles simply, overlooking how individual components (the model, the benchmark items, and the rating scale) impact the final results.
为了解决这一局限性,本文将拉施测量理论(Rasch Measurement Theory, RMT)引入到大模型评估中。RMT 是一种心理计量学框架,旨在将有序评分(ordinal ratings)分解为公共量表上可分离的各个维度,同时提供稳健的诊断工具来识别评委偏差和校准失当的测量结果。
To address this limitation, this paper introduces Rasch Measurement Theory (RMT) into LLM evaluation. RMT is a psychometric framework designed to decompose ordinal ratings into separable facets on a common scale while providing robust diagnostics to identify rater biases and miscalibrated measurements.
以《测量仇恨言论》(Measuring Hate Speech)语料库(一个本质上基于 RMT 构建的数据集)为案例研究,作者评估了来自各种模型家族和能力层级的九个大模型的标注表现。研究结果表明,大模型在以下几个关键领域系统性地偏离了人类评委: * 严重程度(Severity) * 题目级校准(Item-level calibration) * 题目顺序鲁棒性(Question-order robustness) * 目标群体敏感性(Target-identity sensitivity) * 评分量表利用率(Rating scale utilization)
Using the Measuring Hate Speech corpus as a case study—a dataset built inherently on RMT—the authors evaluate annotations from nine LLMs across various families and capability tiers. The findings reveal that LLMs systematically deviate from human raters in several key areas: * Severity * Item-level calibration * Question-order robustness * Target-identity sensitivity * Rating scale utilization
作者认为,标准的评估实践掩盖了这些关键偏差,并主张在“大模型作为考生”、“大模型作为裁判”和“大模型作为评委”的评估范式中全面整合 RMT。
The authors argue that standard evaluation practices obscure these critical biases and advocate for the integration of RMT across LLM-as-examinee, -judge, and -rater evaluation paradigms.
Abstract
大模型现在处于评估的各个环节:它们是基准测试中的受试考生、其他模型输出的裁判,以及人类生成内容的评委。每种范式都可以被视为一个测量问题,即通过评委使用测量工具(如基准测试)的项目来探测对象的潜在属性。标准的评估实践往往忽略了每个核心组件对最终结果的贡献,从而限制了我们对所测量内容的理解。
LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters. Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured.
拉施测量理论(RMT)非常适合解决这类问题。RMT 将有序评分分解为公共量表上可分离的各个维度。它还提供了一系列诊断工具,可以识别校准失当的测量结果和评委偏差。
Rasch measurement theory (RMT) is well-suited to this kind of problem. RMT decomposes ordinal ratings into separable facets on a common scale. It further provides a battery of diagnostics that can identify miscalibrated measurements and rater biases.
我们展示了一个将 RMT 应用于“大模型作为评委”范式的案例研究,使用的是《测量仇恨言论》语料库,该语料库的构建结构本身就是在 RMT 下完成的。我们将一系列多维拉施模型(many-facet Rasch models)拟合到跨越多个模型家族和能力水平的九个大模型的标注中。我们的分析表明,大模型在严重程度、题目级校准、题目顺序鲁棒性、目标群体敏感性和评分量表使用方面与人类评委存在系统性差异,而这些差异在标准的评估实践中都会被掩盖。总的来说,我们认为 RMT 应当纳入评估大模型作为考生、裁判和评委范式的工具箱中。
We present a case study of RMT applied to the LLM-as-rater paradigm using the Measuring Hate Speech corpus, whose construct was itself built under RMT. We fit a series of many-facet Rasch models to annotations from nine LLMs spanning families and capability levels. Our analyses show that LLMs systematically differ from human raters in severity, item-level calibration, question-order robustness, target-identity sensitivity, and rating scale use, which all would be obscured by standard evaluation practice. Overall, we argue that RMT belongs in the toolkit for evaluating LLM-as-examinee, -judge, and -rater paradigms.
Article Assets & External Links
- 全文选项:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- 引用与参考:
- NASA ADS
- Google Scholar
- Semantic Scholar
- Full-Text Options:
- View PDF
- HTML Version (Experimental)
- TeX Source
- Citations & References:
- NASA ADS
- Google Scholar
- Semantic Scholar
