文章背景与核心概要
本文探讨了在使用有界且存在截断(censored)特性的评分量表对大语言模型(LLM)裁判进行审计时所面临的方法论缺陷。作者通过理论推导与实证检验表明,在此类量表上应用双重差分(Difference-in-Differences, DiD)设计,在数学上可能会制造出虚假的交互作用效应,从而误导研究结论。
在一项针对教育学大模型裁判的预注册审计中,研究人员发现其主要的预注册终点(评估学习者画像对支架式教学偏好的影响)呈现零效应(\(+0.085\) 分,\(p = 0.684\))。然而,一个名义上显著的交互作用效应(\(+0.378\),\(p = 0.002\))却意外出现——这纯粹是由量表底层截断(floor censorship)以及观察到的严重程度偏移(severity shifts)所导致的统计假象,而非真正的差异化偏好。作者通过闭式解推导了这一失效机制,并证明其可以直接从标准的审计数据中进行测量。
Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
arXiv: 2608.27309 [cs.CL]
Submitted: August 27, 2026
Authors: Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li, Hongyang Zhang
📌 Summary
本文研究了在使用有界且截断的评分量表审计大语言模型(LLM)裁判时的方法论缺陷。作者证明,在这些量表上采用双重差分(DiD)设计,会在数学上制造出虚假的交互作用效应。
通过对一个教育学裁判进行预注册审计,主要注册终点(评估学习者画像对脚手架偏好的影响)为零效应(\(+0.085\) 分,\(p = 0.684\))。然而,一个名义上显著的交互效应(\(+0.378\),\(p = 0.002\))却浮现出来,这纯粹是量表底限截断和观察到的严重程度偏移的产物,而不是真正的差异化偏好。作者用闭式解推导了这一失效机制,证明它可以直接从标准的审计数据中进行测量。
This paper investigates methodological flaws in auditing Large Language Model (LLM) judges using bounded, censored rating scales. The authors demonstrate that employing a Difference-in-Differences (DiD) design on such scales can mathematically manufacture illusory interaction effects.
Through a pre-registered audit of a pedagogy judge, the primary registered endpoint (evaluating the effect of a learner profile on scaffolding preference) was null (\(+0.085\) points, \(p = 0.684\)). However, a nominally significant interaction effect (\(+0.378\), \(p = 0.002\)) emerged purely as an artifact of scale floor censorship and observed severity shifts, rather than true differential preference. The authors derive this failure mechanism in closed form, proving it is directly measurable from standard audit data.
📖 Abstract
LLM裁判的审计通过对比匹配条件来验证某种偏见,其中最强的设计会进行两次差分:首先在两个候选响应之间进行条目内对比,然后在被操纵的属性之间再次差分,最后从一个有界的评分量表中读出结果。我们表明,这个终点并没有在报告它的量表上得到识别。
双重差分的每一项都受到其自身份额的截断,因此观察到的统计量将差异化偏好与差异化衰减混淆了:当两个响应受到不均匀的截断时,对两类响应共同起作用的严重程度偏移就会制造出一个交互作用——正因为优秀的刺激物恰好位于距离边界不均等的位置。
我们在一个冻结的教育学裁判的预注册审计中展示了这种失效,该审计在其 990 次调用之前就已经完成密封。注册的主要终点——陈述的学习者画像对裁判脚手架偏好的影响——为零:\(+0.085\) 分(95% BCa \([-0.167, +0.353]\),\(p = 0.684\))。审计中唯一名义上显著的交互作用 \(+0.378\)(\(p = 0.002\))并未被识别为偏好:一个完全不包含差异化偏好的结构,仅从观察到的严重程度偏移和量表底限本身,就重现了其中 79% 到 85% 的效果。我们以闭式解推导了该机制,并表明其贡献可以从审计自身的评分中进行测量。
Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it.
Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them.
We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: \(+0.085\) points (95% BCa \([-0.167, +0.353]\), \(p = 0.684\)). The audit's one nominally significant interaction, \(+0.378\) (\(p = 0.002\)), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85% of it from the observed severity shift and the scale floor alone. We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings.
📋 Metadata & Article Details
- 学科主题: 计算与语言 (
cs.CL);人工智能 (cs.AI);计算机与社会 (cs.CY)- MSC 分类: 68T50(主要);68T05、97U50、62G10(次要)
- ACM 分类: K.3.1;I.2.7
- DOI: 10.48550/arXiv.2608.27309
- 许可协议: 知识共享 署名-非商业性使用-禁止演绎 4.0 国际版 (查看下方许可图标)
- Subjects: Computation and Language (
cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY) - MSC Classes: 68T50 (Primary); 68T05, 97U50, 62G10 (Secondary)
- ACM Classes: K.3.1; I.2.7
- DOI: 10.48550/arXiv.2608.27309
- License: Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (View license icon below)

🔗 Links & Resources
- 全文访问: 查看 PDF | HTML(实验性) | TeX 源码
- 引用与文献目录: Google Scholar | Semantic Scholar | NASA ADS
- Full-Text Access: View PDF | HTML (Experimental) | TeX Source
- Citations & Bibliography: Google Scholar | Semantic Scholar | NASA ADS