文章背景与核心概要
当前的“大模型作为裁判”(LLM-as-a-judge)系统评估通常依赖于一致性指标以及对表面级扰动的鲁棒性。然而,高可靠性并不等同于真正的构念效度。本文通过一个二维画像对评估者的构念效度进行了形式化定义:1. 不变性(\(S\)):在保持构念不变的编辑下,裁决保持不变的概率;2. 构念敏感性(\(R\)):在最小限度改变构念的编辑下,裁决发生改变的概率。
作者证明了 \(S\) 和 \(R\) 是独立的指标,这意味着单一的标量总结无法捕捉所有关键的比较。在对 7 个裁判和 4 个领域进行的评估中——使用了 7 种改变构念的干预类型和 5 种仅改变文体(register-only)的控制组——研究结果揭示了当前模型的重大局限性。即使在匹配的不变性阈值 \(S \ge 0.90\) 下,裁判的平均不变性很高(\(S = 0.945\)),但敏感性却异常低(\(R = 0.319\))。此外,敏感性在范围(scope)和强度(strength)编辑之间存在差异(\(R_{\text{scope}} = 0.383\) 对比 \(R_{\text{strength}} = 0.262\)),在所有 7 个裁判中表现出一致的 \(+0.121\) 的差距。对公共标签集的审计还显示,表面级预测器可以在配对模式下复制 55%–67% 的标签。归根结底,该研究强调,高裁判一致性可能会掩盖对被评估构念实际变化的弱敏感性,并主张联合报告不变性和敏感性,同时进行严格的验证集审计。
A Judge Should Know What Changed: Construct Validity for LLM-as-a-Judge Evaluation
Authors: Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
Subject: Artificial Intelligence (cs.AI)
arXiv ID: 2608.24419
Submission Date: August 25, 2026
Summary
当前 evaluations 的 LLM-as-a-judge 系统通常依赖于一致性 metrics 和针对表层扰动的 robustness。然而,高 reliability 并不等同于真正的 construct validity。
Current evaluations of LLM-as-a-judge systems typically rely on agreement metrics and robustness against surface-level perturbations. However, high reliability does not equate to genuine construct validity.
本文通过一个二维 profile 将评估者的 construct validity 进行了形式化定义: 1. 不变性(Invariance,\(S\)): 在 construct-preserving edits 下,verdict 保持不变的概率。 2. 构念敏感性(Construct Sensitivity,\(R\)): 在 minimal construct-changing edits 下,verdict 发生改变的概率。
This paper formalizes construct validity for evaluators through a two-dimensional profile: 1. Invariance (\(S\)): The probability that a verdict remains unchanged under construct-preserving edits. 2. Construct Sensitivity (\(R\)): The probability that a verdict changes under minimal construct-changing edits.
作者证明了 \(S\) 和 \(R\) 是独立的 metrics,这意味着没有单一的 scalar summary 可以捕捉所有关键的比较。在对 7 个 judges 和 4 个 domains 的 evaluation 中——使用了 7 种 construct-changing intervention types 和 5 种 register-only controls——findings 揭示了当前 models 的显著 limitations。即使在匹配的 invariance threshold \(S \ge 0.90\) 下,judges 的平均 invariance 很高(\(S = 0.945\)),但 sensitivity 却非常低(\(R = 0.319\))。此外,sensitivity 在 scope 和 strength edits 之间有所不同(\(R_{\text{scope}} = 0.383\) 与 \(R_{\text{strength}} = 0.262\)),在所有 7 个 judges 中表现出一致的 \(+0.121\) 的 gap。Auditing public label sets 还表明,surface-level predictors 可以在 paired modes 下复制 55%–67% 的 labels。最终,该研究强调,高 judge agreement 可能会掩盖对 evaluated construct 实际变化的微弱 sensitivity,并倡导联合报告 invariance 和 sensitivity,同时进行严谨的 validation-set audits。
The authors demonstrate that \(S\) and \(R\) are independent metrics, meaning no single scalar summary can capture all critical comparisons. Across an evaluation of 7 judges and 4 domains—using 7 construct-changing intervention types and 5 register-only controls—the findings reveal significant limitations in current models. Even at a matched invariance threshold of \(S \ge 0.90\), judges averaged high invariance (\(S = 0.945\)) but remarkably low sensitivity (\(R = 0.319\)). Furthermore, sensitivity varied between scope and strength edits (\(R_{\text{scope}} = 0.383\) versus \(R_{\text{strength}} = 0.262\)), showing a consistent \(+0.121\) gap across all 7 judges. Auditing public label sets also revealed that surface-level predictors can replicate 55%–67% of labels in paired modes. Ultimately, the research emphasizes that high judge agreement can mask a weak sensitivity to actual changes in the evaluated construct, advocating for the joint reporting of invariance and sensitivity alongside rigorous validation-set audits.
Article Metadata
- Comments: 39 pages, 10 figures, 11 tables
- DOI: 10.48550/arXiv.2608.24419
- License: Creative Commons Attribution 4.0 International
