文章背景与核心概要
当前的AI对齐范式——人类反馈强化学习(RLHF)——在处理人类多元化偏好时,往往缺乏严格的社会选择学保证。本文提出了一种根本性的范式转变:通过将对齐问题重新表述为凸影响空间上的线性优化,作者成功弥合了AI开发与福利经济学、机制设计等成熟领域之间的鸿沟。
这种方法使研究人员能够将复杂的社会期望(例如对群体伤害的约束)直接转化为可执行的对齐协议,从而为协调冲突的人类价值观提供了一个坚实的数学基础。该研究不仅阐明了对齐协议如何转化为福利后果,还通过实证展示了其在肾脏分配、慈善食品分发、大模型响应和电车难题等真实人类偏好场景中的应用价值。
Algorithilc Impact Reveals the Hidden Social Choice Structure of Alignment
Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment
Authors: Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia
Date: August 25, 2026
Subject: Artificial Intelligence (cs.AI)
arXiv ID: 2608.24046
摘要
当AI算法做出影响多人的决策时,对其进行对齐就变成了一个社会选择问题:如何将人们对系统行为的不同偏好进行调和,并将其聚合成一个单一连贯的模型?
对齐前沿AI模型的标准方法——人类反馈强化学习(RLHF),在很大程度上避开了这个问题,并且其社会选择保证较差。然而,目前尚不清楚应该用什么替代方案来取代它。我们表明,通过直接关注算法的福利后果,可以将对齐问题重新表述为凸影响空间上的线性优化,这使得它能够适用于福利经济学和机制设计的标准工具集。
这种重新表述阐明了对齐协议如何转化为福利后果,以及反过来,社会规划者对福利后果的期望约束如何能够被转化回对齐协议。我们应用这种转化来证明,按议题投票(voting-by-issues)和随机独裁(random-dictatorship)机制具有策略防范性(strategyproof)和一致性(unanimous)。为了展示反向路径,我们还利用影响表示法推导出一类对齐协议,它们在满足各种社会期望(如对个人或群体伤害的约束)的前提下最大化功利主义社会福利。我们通过使用关于肾脏分配、慈善食品分发、大模型响应和电车难题的真实人类偏好,从实证角度说明了这些对齐协议的福利含义。
Abstract
When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of social choice: how should people's divergent preferences about system behavior be reconciled and aggregated into a single coherent model?
The standard approach to aligning frontier AI models—reinforcement learning from human feedback—largely sidesteps this question and has poor social choice guarantees. However, it remains unclear what alternative should replace it. We show that, by focusing directly on an algorithm's welfare consequences, the alignment problem can be reformulated as linear optimization over a convex impact space, which makes it amenable to the standard toolkit of welfare economics and mechanism design.
This reformulation clarifies how alignment protocols translate into welfare consequences and, conversely, how a social planner's desired constraints on welfare consequences can be translated back into alignment protocols. We apply this transformation to show that voting-by-issues and random-dictatorship mechanisms are strategyproof and unanimous. Demonstrating the reverse direction, we also apply the impact representation to derive a family of alignment protocols that maximize utilitarian social welfare subject to various social desiderata, such as bounds on individual or group harm. We illustrate the welfare implications of these alignment protocols empirically using real human preferences over kidney allocation, charitable food distribution, LLM responses, and trolley problems.
访问与资源
Access & Resources
引用
如果您发现这项研究有用,可以通过以下方式对其进行引用:
Citation
If you find this research useful, you can cite it via: * NASA ADS * Google Scholar * Semantic Scholar