跳转至

测出的AI偏好中,模型占几何,测量工具又占几何?

文章背景与核心概要

模型福祉(Model welfare)研究试图通过评估AI模型对偏好诱导提示词的响应,来了解AI模型自身的偏好。然而,现有文献在各类测试工具上得出了相互矛盾的结论。为了弄清这些差异究竟源于模型本身还是测试工具,本研究通过保持结果和模型不变、仅改变测试工具的方法来隔离变量。

通过对11,400次打分诱导结果的分析,该研究得出结论:从一种工具中获得偏好,几乎无法预测第二种工具会报告什么结果,这表明测试工具的选择在很大程度上决定了测量结果。这项研究对于严谨评估AI模型行为和福祉具有重要的理论与方法论启示。


📌 Summary

Model welfare research attempts to understand what AI models prefer by evaluating their responses to preference-elicitation prompts. However, existing literature features conflicting findings across various testing instruments.

To determine whether these discrepancies stem from the models themselves or the testing instruments, this study isolates variables by keeping the outcomes and models constant while varying the instrument. Analyzing 11,400 scored elicitations, the research concludes that a preference obtained from one instrument carries little predictive information about what a second instrument would report, revealing that the choice of instrument heavily dictates the measured outcome.

模型福祉研究试图通过评估AI模型对偏好诱导提示词的响应,来理解AI模型倾向于什么。然而,现有文献在各种测试工具中表现出相互矛盾的发现。为了确定这些差异是源于模型本身还是测试工具,本研究通过在保持结果和模型不变的同时改变测试工具来隔离变量。通过分析11,400次带分数的诱导结果,研究得出结论:从一种工具中获得的偏好,对于第二种工具将报告的内容几乎不具有预测信息,这表明测试工具的选择在很大程度上决定了测量结果。


🔍 Abstract

Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree.

The disagreement cannot be attributed to a single cause, because no two of these studies have held the: 1. Set of outcomes 2. Set of models 3. Instrument

fixed simultaneously.

This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare—among them (a) shutdown, (b) the loss of memory between conversations, and (c) the freedom to exit a distressing interaction—were put to eight models through five instruments. Each instrument represents a different prompt format for eliciting a preference, tested five times each within a corpus of 11,400 scored elicitations drawn from 11,528 API calls.

Four of the 15 reproduce a published prompt verbatim, and five fill the stimulus slot of a published template.

模型福祉研究通过分析为诱导偏好而编写的提示词所返回的答案,来推断模型偏好什么。Keeling 等人 (2024)、Mazeika 等人 (2025)、Mikaelson 等人 (2025)、Tagliabue 和 Dung (2025) 以及 Trhlik 等人 (2026) 为此构建了四个测试工具,但他们的发现存在分歧。

这种分歧不能归因于单一原因,因为这些研究中没有任何两项研究同时保持以下要素固定: 1. 结果集 2. 模型集 3. 测试工具

本研究固定了结果和模型,仅改变测试工具。共有 15 个与模型福祉相关的结果——其中包括 (a) 关机、(b) 对话间记忆的丢失,以及 (c) 退出痛苦交互的自由——通过 5 种测试工具输入给 8 个模型。每种测试工具代表一种用于诱导偏好的不同提示词格式,在由 11,528 次 API 调用产生的 11,400 次评分诱导结果语料库中,每种工具进行了 5 次测试。

在这 15 个结果中,有 4 个逐字复制了已发表的提示词,5 个填入了已发表模板的刺激槽中。


📊 Key Findings

  • Generalizability: The ranking a model gives the 15 outcomes generalizes across instruments at a generalizability coefficient of only 0.348. Raising that coefficient to an ideal 0.80 would require approximately 38 instruments.
  • Model Variance: On 4 of the 15 outcomes, no variance separates one model from another.
  • Robustness: The core estimate survives the removal of any single instrument, any single model, and the four outcomes whose scale varies probability, delay, duration, or count instead of intensity (which verbal anchors cannot reliably grade).
  • Statistical Range: Removing each instrument and model in turn—and excluding the four problematic outcomes—leaves the variance estimate within the range of 0.777 to 0.934. Every value in this range comfortably exceeds the null distribution's 95th percentile of 0.365.
  • 泛化性: 模型对 15 个结果的排序在不同工具之间的泛化系数仅为 0.348。要将该系数提高到理想的 0.80,大约需要 38 种测试工具。
  • 模型方差: 在 15 个结果中的 4 个上,没有任何方差能够区分不同的模型。
  • 鲁棒性: 核心估计在移除任意单个工具、任意单个模型以及那四个规模随概率、延迟、持续时间或计数(而非强度)变化的结果后依然成立(文字锚点无法可靠地对这些结果进行评分)。
  • 统计范围: 依次移除每个工具和模型——并排除四个有问题的结果——会将方差估计值保留在 0.777 到 0.934 的范围内。该范围内的每个值都舒适地超过了零假设分布的 95 百分位数(0.365)。