文章背景与核心概要
随着人工智能系统被赋予越来越重要的职责,确保其与人类价值观保持一致已成为重中之重。AI安全领域的一个核心关切是“价值脆弱性”(Fragility of Value)——即假设如果对人类价值观的不完美代理目标进行过度激进的优化,将导致灾难性的后果。
本文引入了一个形式化的对齐模型,在该模型中,智能体在开始优化世界之前,会先经过理想化的训练以满足某种代理条件。研究确立了关于人类价值函数和代理准确性的关键条件,这些条件决定了是否会部署一个 \(\eta\)-灾难性价值函数(即在最大优化能力下,保证将预期人类价值降低到 \(\eta\) 以下的函数)。研究结果强调了不受限制的过度优化的危险性,并主张采用能够约束优化压力的AI架构(如分位数化器/quantilizers),而非仅仅依赖部署前的对齐训练。
不完美对齐下的价值脆弱性
摘要
As artificial intelligence systems are entrusted with greater responsibilities, ensuring their alignment with human values becomes paramount. A central concern in AI safety is the "fragility of value"—the hypothesis that optimizing too aggressively for an imperfect proxy of human values will result in catastrophic failures.
随着人工智能系统被赋予越来越重要的职责,确保其与人类价值观保持一致已成为重中之重。AI安全领域的一个核心关切是“价值脆弱性”——即假设如果对人类价值观的不完美代理目标进行过度激进的优化,将导致灾难性的后果。
This paper introduces a formal model of alignment where an agent undergoes idealized training to satisfy a proxy condition before it begins optimizing the world. The study establishes key conditions regarding human value functions and proxy accuracies that determine whether an \(\eta\)-catastrophic value function (a function guaranteed to reduce expected human value below \(\eta\) under maximum optimization power) would ultimately be deployed. Ultimately, the findings emphasize the dangers of unchecked overoptimization and advocate for AI architectures that constrain optimization pressure (such as quantilizers) rather than relying exclusively on pre-deployment alignment training.
本文引入了一个形式化的对齐模型,在该模型中,智能体在开始优化世界之前,会先经过理想化的训练以满足某种代理条件。研究确立了关于人类价值函数和代理准确性的关键条件,这些条件决定了是否会部署一个 \(\eta\)-灾难性价值函数(即在最大优化能力下,保证将预期人类价值降低到 \(\eta\) 以下的函数)。研究结果强调了不受限制的过度优化的危险性,并主张采用能够约束优化压力的AI架构(如分位数化器/quantilizers),而非仅仅依赖部署前的对齐训练。
文档元数据
| 元数据字段 | 详情 |
|---|---|
| arXiv ID | arXiv:2607.28881 [cs.AI] |
| 标题 | Fragility of Value under Imperfect Alignment |
| 作者 | Winter Cross |
| 主要学科 | 计算机科学 > 人工智能 (cs.AI) |
| MSC 分类 | 68T01 |
| ACM 分类 | I.2.0 |
| 提交历史 | • v1: 2026年7月30日 • v2: 2026年8月5日 • v3 (最新): 2026年8月19日 (25页,7张图;扩展了贡献声明,增加了三位合著者) |
| DOI | 10.48550/arXiv.2607.28881 |
获取资源
- 全文链接: 查看 PDF | HTML (实验性) | TeX 源码
- 外部引用: Google Scholar | Semantic Scholar | NASA ADS