跳转至

语言模型认为谁更有能力?职业偏见的机械可解释性分析

文章背景与核心概要

尽管现代语言模型(LMs)经常能通过各种行为层面的偏见评估,但它们是真正消除了潜在的偏见关联,还是仅仅学会了压制其表露,这依然是一个未解之谜。本文引入了一个因果框架,通过检查模型对用户能力的内部表征及其可观察的输出这两个不同的测量点,来深入研究职业偏见。

通过推导用户专业知识表征的操纵向量(steering vectors),作者证明了这些内部表征在问答和招聘任务中都对模型行为起到了因果中介的作用。对开源模型的评估表明,人口统计学属性(如性别、种族和社会经济地位)会不知不觉地塑造模型对专业知识的内部评估——即使在标准行为指标未检测到任何差异的情况下也是如此。归根结底,这些隐藏的表征在受到干预时会驱动偏见的下游行为,这凸显了仅靠行为指标无法捕捉的关键AI安全与评估失效模式。


Summary

While modern language models (LMs) frequently pass behavioral bias evaluations, it remains an open question whether they have truly eliminated underlying bias associations or simply learned to suppress their expression. This paper introduces a causal framework to investigate occupational bias by examining it at two distinct measurement points: the model's internal representation of a user's competence and its observable outputs. By deriving steering vectors for user expertise representations, the authors demonstrate that these internal representations causally mediate model behavior in both question-answering and hiring tasks. Evaluating open-weight models reveals that demographic attributes (such as gender, race, and socioeconomic status) inadvertently shape a model's internal assessment of expertise—even when standard behavioral metrics register no disparity. Ultimately, these hidden representations can drive biased downstream behavior under intervention, highlighting critical AI safety and evaluation failure modes that behavioral metrics alone fail to capture.


Paper Metadata

  • arXiv Identifier: arXiv:2608.20347 [cs.CL]
  • DOI: 10.48550/arXiv.2608.20347
  • Primary Subject: Computation and Language (cs.CL)
  • Secondary Subjects: Artificial Intelligence (cs.AI), Computers and Society (cs.CY)
  • Authors: Keren Fuentes, Aaron Mueller
  • Submitted: June 15, 2026

Abstract

语言模型(LMs)经常能够通过行为偏见评估,但目前尚不清楚它们是彻底不再表征引发偏见的底层关联,还是仅仅学会了不将其表现出来。在这项研究中,我们表明,即使在行为偏见不可见的情况下,表征偏见通常也是可以检测到的。我们引入了一个因果框架,将职业偏见分解为两个测量点:模型对用户能力的内部表征,以及其可观察到的输出。我们推导出了用户专业知识表征的操纵向量,并验证了它们在问答任务和招聘任务中对模型行为具有因果中介作用。将该框架应用于多个开源权重模型后,我们发现人口统计学属性(如性别、种族和社会经济地位)会影响模型对用户专业知识的表征,即使在行为指标未检测到不同人口群体间存在差异的情况下也是如此。我们证明,在干预下,这些模型表征可以影响下游行为,这表明单纯依靠行为指标可能无法检测到某些失效模式。

Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that representational biases are often detectable, even when behavioral biases are not visible. We introduce a causal framework that decomposes occupational bias into two measurement points: a model's internal representation of a user's competence, and its observable outputs. We derive steering vectors for representations of user expertise, and verify that they causally mediate model behavior in both a question-answering task and a hiring task. Applying this framework to several open-weight models, we find that demographic attributes, such as gender, race, and socioeconomic status, influence a model's representation of user expertise, even in cases where behavioral metrics detect no disparity between demographics. We show that these model representations can influence downstream behavior under intervention, suggesting failure modes that behavioral metrics alone may not detect.


访问链接与资源: