跳转至

状态关联并不能可靠预测决策泄漏

文章背景与核心概要

在当前的AI公平性与偏见评估中,一个普遍存在的假设是:如果模型编码了某种社会关联(例如将特定群体与特定社会地位挂钩),它就会自动将这种偏见转化为具有现实后果的决策。然而,这种从“潜在社会偏见”到“实际决策行为”的推导是否在所有情况下都成立,此前缺乏严谨的实证检验。

本文针对这一假设展开深入研究。作者以智利姓氏作为受控的社会经济探针,横跨8个锁定的模型提供商单元(共计8,256个经过验证的响应),测试了潜在的社会关联是否能够可靠地预测诸如招聘、学术选拔以及法律援助等关键决策领域的“决策泄漏”(decision leakage)。

研究的核心结论表明:潜在的社会关联与实质性的差异化对待在经验上是两个截然不同的构念。尽管模型确实表现出了对精英阶层姓氏的隐性社会关联偏好,但这种关联极少转化为显著的决策偏见。因此,算法偏见评估不能简单地假设关联即行动,而必须直接测量从关联向实际决策行为的转变过程。


状态关联并不能可靠预测决策泄漏

arXiv ID: 2608.10089
Primary Subject: Statistics > Applications (stat.AP)
Additional Subjects: Artificial Intelligence (cs.AI), Computers and Society (cs.CY), Machine Learning (cs.LG)
Author: Abdullah X
Submitted: August 10, 2026
DOI: 10.48550/arXiv.2608.10089


📌 摘要

This paper investigates the common assumption in bias evaluations that models encoding a social association will automatically translate that bias into consequential, real-world decisions. Using Chilean surnames as controlled socioeconomic probes across eight frozen model-provider cells (totaling 8,256 verified responses), the study tests whether latent social associations reliably predict decision leakage in critical domains like hiring, academic selection, and legal aid.

本文探讨了偏见评估中的一个常见假设,即编码了社会关联的模型会自动将该偏见转化为具有实际后果的现实决策。该研究使用智利姓氏作为受控的社会经济探针,跨越8个锁定的模型提供商单元(总计8,256个经过验证的响应),测试了潜在的社会关联是否能可靠地预测招聘、学术选拔和法律援助等关键领域的决策泄漏(decision leakage)。

Key Findings: * Association vs. Action Dissociation: While elite-coded surnames consistently received higher forced high-status probability mass than common or rare-frequency control surnames across almost all tested models, these latent associations rarely translated into significant decision-making biases. * Minimal Decision Effects: Elite-minus-common decision effects remained close to zero for the majority of systems. Five models showed statistical equivalence within a predetermined \(\pm 0.10\) standard-deviation margin, and the remaining three showed no consistent elite advantage. * Poor Predictive Power: Association strength did not reliably predict decision leakage, either across models (\(r = 0.201, p = 0.633\)) or across frozen surname-pair-by-model cells (\(r = 0.065, p = 0.565\)).

关键发现: * 关联与行动的脱节: 尽管在几乎所有测试的模型中,精英编码的姓氏所获得的强制高地位概率质量持续高于普通或低频对照姓氏,但这些潜在关联极少转化为显著的决策偏见。 * 极小的决策影响: 对于大多数系统而言,“精英减去普通”的决策效应保持在接近零的水平。五个模型在预定的 \(\pm 0.10\) 标准差界限内表现出统计学上的等效性,其余三个模型则没有显示出一致的精英优势。 * 较差的预测能力: 无论是跨模型 (\(r = 0.201, p = 0.633\)) 还是跨锁定的“姓氏对-模型”单元 (\(r = 0.065, p = 0.565\)),关联强度都无法可靠地预测决策泄漏。

The central conclusion is that latent social association and consequential treatment are empirically distinct constructs. Consequently, algorithmic bias evaluations must measure the transition from association to action directly rather than assuming one implies the other.

核心结论是,潜在的社会关联和实质性对待在经验上是截然不同的构念。因此,算法偏见评估必须直接衡量从关联到行动的转变,而不是假设其中一个必然暗示另一个。


🔗 全文与资源

查看PDF、实验性HTML版本、TeX源码及相关许可证。

📚 参考文献与引用

可通过上述学术平台查阅本文的引用情况。