文章背景与核心概要
随着大语言模型(LLM)越来越多地被部署为评估器(即“LLM-as-a-judge”),人们对自我偏好(即倾向于青睐自身输出的倾向)的担忧日益增加。然而,以往的研究主要集中在生成的文本上,此时文体特征与回答质量密不可分,这使得很难将真正的自我偏好与其它混杂因素剥离。为了解决这一局限性,研究人员将评估对象从生成的文本转向了叙事约束选择——这类选择不带有模型特定的文体足迹,同时保留了可恢复的模型特定签名。通过涉及十个LLM的两项对照实验,该研究揭示了在盲评和匹配质量条件下,模型归因标签如何显著扭曲评估结果。
该研究的两大核心贡献在于:1)证明了署名归因是评估偏差的一个独立驱动因素;2)确立了开放式、无标准答案的任务可以作为研究LLM裁判行为的受控工具。
Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
- arXiv ID: 2608.18091 [cs.CL]
- Subjects: Computation and Language (
cs.CL); Artificial Intelligence (cs.AI) - Authors: Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
- Submitted: June 6, 2026
- DOI: 10.48550/arXiv.2608.18091
Summary
As large language models (LLMs) are increasingly deployed as evaluators ("LLM-as-a-judge"), concerns regarding self-preference (the tendency to favor one's own outputs) have grown. However, previous studies focused primarily on generated text, where stylistic features and response quality are inextricably linked, making it difficult to isolate true self-preference from other confounding factors.
To resolve this limitation, the researchers shifted the evaluation target from generated text to narrative constraint selections, which carry no model-specific stylistic footprint while retaining a recoverable model-specific signature. Through two controlled experiments involving ten LLMs, the study reveals:
- Under Blind Evaluation: Once selection quality and evaluator severity are controlled, self-preference largely disappears across three out of four rubric dimensions, and actually reverses on the fourth (where judges rate their own selections as less original).
- Under Matched Quality: Even without naming any specific model, self- and other-labels alone induce a bidirectional bias: LLM judges systematically inflate scores for self-labeled selections and deflate scores for other-labeled selections, regardless of the actual source.
Key Contributions
- Authorship Attribution as a Bias Driver: Demonstrates that mere attribution labels significantly skew evaluation outcomes.
- Controlled Evaluation Instruments: Establishes that open-ended, ground-truth-free tasks can effectively serve as controlled frameworks for studying LLM judge behaviors.
随着大语言模型(LLM)越来越多地被部署为评估器(“LLM-as-a-judge”),LLM中的自我偏好(即偏爱自身输出的倾向)引发了人们对评估可靠性的日益担忧。然而,以往的研究主要集中在生成的文本上,在这些文本中,文体特征和回答质量不可避免地交织在一起。因此,现有的测量方法无法将真正的自我偏好与这些混杂因素区分开来。我们通过改变评估对象来解决这一问题:十个LLM不再评估生成的文本,而是评估叙事约束选择(narrative constraint selections),这些选择不带有模型特定的文体指纹,但保留了可恢复的模型特定签名。我们进行了两项得出了不同发现的实验。在盲评下,一旦控制了选择质量和评估器的严格程度,自我偏好在很大程度上就会消失。它在四个评分维度的三个维度上消失了,并在第四个维度上逆转——裁判认为他们自己的选择原创性较低。然而,在质量匹配的情况下,仅凭自我标签和他人标签——在不命名任何模型的情况下——就会导致双向偏差:无论选择的实际来源如何,LLM裁判都会系统性地提高带有自我标签的选择的分数,并降低带有他人标签的选择的分数。我们作出了两项贡献:1)署名归因是评估偏差的一个独特驱动因素,2)开放式、无真值的任务可以作为研究LLM裁判行为的受控工具。
Abstract
As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this by changing the object of evaluation: instead of judging generated text, ten LLMs assess narrative constraint selections, which carry no model-specific stylistic fingerprint yet retain a recoverable model-specific signature. We run two experiments that yield distinct findings. Under blind evaluation, self-preference largely disappears once selection quality and evaluator severity are controlled. It vanishes on three of four rubric dimensions and reverses on the fourth, where judges rate their own selections as less original. Under matched quality, however, self- and other-labels alone -- without naming any model -- shift scores bidirectional: LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection's actual source. We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) open-ended, ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.
