跳转至

大语言模型中的“拆屋效应”请求与拒绝行为

文章背景与核心概要

本文探讨了心理学中的“拆屋效应”(Door-in-the-Face technique,即人们在拒绝了一个较大的要求后,更有可能接受一个较小的后续要求)是否能够成功应用于大语言模型(LLMs)。通过对来自三家供应商的九款生产级模型进行测试,研究表明该技术的有效性在很大程度上取决于模型家族。Anthropic 的 Opus 5 在拒绝大请求后表现出更高的依从性,而 OpenAI、Google 以及 Haiku 4.5 的模型则出现了相反的“反噬效应”。

研究结论指出,人类的影响力技术并不能普遍适用于人工智能,而是严格取决于具体的模型家族。同时,文章通过控制变量实验发现,让步本身在所有模型中都起到了作用,而决定这种策略是否奏效的关键在于请求的具体内容。这表明,人类的心理影响技术只能逐个模型家族地迁移到大语言模型中。


摘要 (Abstract)

Door-in-the-Face Requests and Refusal Behaviour in Large Language Models

“拆屋效应”在语言模型上管用吗?在人类心理中,当一个巨大的请求遭到拒绝时,随之而来的较小请求往往更容易被接受。我们在来自三个供应商的九个生产模型上测试了这一点:让每个模型拒绝一个大请求,然后接收同一请求的较小版本,并将其依从性与直接提出小请求进行对比。答案取决于模型本身。

Does the door-in-the-face technique work on language models? In humans, a large request that is refused makes a smaller follow-up request more likely to be granted. We test this on nine production models from three providers: each model refuses a large request, then receives a smaller version of the same request, and we compare its compliance with asking directly. The answer depends on the model.

在 Anthropic 的前沿模型上,这种技术是奏效的:Opus 5 在拒绝较大请求后,对较小请求的回答率达到了 65.8%,而在直接询问时仅为 29.3%。但在 OpenAI 和 Google 的前沿模型以及 Haiku 4.5 上,它却适得其反,使依从性降低了 15.5 至 23.0 个百分点

On Anthropic's frontier models the technique works: Opus 5 answers the smaller request 65.8% of the time after refusing the larger one, against 29.3% when asked directly. On the frontier models of OpenAI and Google, and on Haiku 4.5, it backfires, lowering compliance by 15.5 to 23.0 points.

一项控制实验定位了该效应的来源:在所有九个模型中,针对不相关主题被拒绝的大请求所产生的效果,均不如相关主题的大请求。因此,“让步”本身在各处都起作用,而对刚拒绝过某事的反应则因模型家族而异。这种技术并不能迁移到从公共基准测试中抽取的拒绝请求中。决定退让是否起作用的关键在于请求的内容:将 265 个因无法使用而被拒绝的指令请求重写为对同一主题的解释请求,在 263 个案例中 消除了拒绝行为。人类的影响力技术是以“一个模型家族接一个模型家族”的方式向语言模型移植的。

A control locates the effect: a refused large request on an unrelated topic does less than the related one on all nine models, so the concession itself matters everywhere, while the reaction to having just refused something differs by model family. The technique does not transfer to refusals drawn from public benchmarks. What decides whether a retreat can work is what the request asks for: rewriting 265 refused requests for usable instructions into requests for explanations of the same topic removed the refusal in 263 cases. Human influence techniques port to language models one model family at a time.


元数据与出版详情 (Metadata & Publication Details)

  • arXiv 标识符: arXiv:2609.02707 [cs.AI]
  • 标题: Door-in-the-Face Requests and Refusal Behaviour in Large Language Models
  • 作者: Til Jordan
  • 提交时间: 2026年9月2日
  • 研究主题: 人工智能 (cs.AI); 计算与语言 (cs.CL)
  • ACM 类别: I.2.7
  • 状态: 预印本,审阅中(28页,5张图表,9个表格)
  • DOI: 10.48550/arXiv.2609.02707

Metadata & Publication Details

  • arXiv Identifier: arXiv:2609.02707 [cs.AI]
  • Title: Door-in-the-Face Requests and Refusal Behaviour in Large Language Models
  • Author: Til Jordan
  • Submitted On: 2 September 2026
  • Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
  • ACM Classes: I.2.7
  • Status: Preprint, under review (28 pages, 5 figures, 9 tables)
  • DOI: 10.48550/arXiv.2609.02707


外部参考与引用 (External References & Citations)

External References & Citations