跳转至

文章背景与核心概要

大语言模型(LLM)在表现出强大能力的同时,也经常展现出各种社会、情感和认知偏差。Eldad Yechiam 和 Adi Tarabeih 的这项最新研究深入探讨了驱动这些偏差的两个关键机制:错误模仿(LLM从逻辑上无关的行为中推断人类偏好)以及对显式偏差人类行为的盲目模仿

通过对经济决策领域的四项研究,作者发现诸如 ChatGPT-4o 和 Qwen 等模型,在面对明显不具备指示性的数据时仍会表现出社会认同偏差,并且在损失厌恶被明确标记为一种偏差时依然会表现出该倾向。值得注意的是,该研究揭示了一个引人深思的现象:记录这些认知偏差的学术论文本身可以充当“自证预言”,直接塑造接触到它们的 LLM 的行为偏差。这项研究不仅揭示了 LLM 的偏差现象,更为理解其背后的底层成分加工过程提供了重要的科学洞察。


Mimicry without understanding: the origins of decision bias in large language models

Summary

Large language models (LLMs) frequently exhibit social, affective, and cognitive biases. This research by Eldad Yechiam and Adi Tarabeih investigates two mechanisms driving these biases: faulty mimicry (where LLMs infer human preferences from logically unrelated behaviors) and mimicry of explicitly biased human behaviors. Across four studies on economic decision-making, models like ChatGPT-4o and Qwen demonstrated social proof biases even when presented with non-indicative data, and displayed loss aversion when it was explicitly framed as a bias. Notably, the study reveals that academic papers documenting these cognitive biases can act as self-fulfilling prophecies, directly shaping the behavioral biases of LLMs exposed to them.

Large language models (LLMs) frequently exhibit social, affective, and cognitive biases. This research by Eldad Yechiam and Adi Tarabeih investigates two mechanisms driving these biases: faulty mimicry (where LLMs infer human preferences from logically unrelated behaviors) and mimicry of explicitly biased human behaviors. Across four studies on economic decision-making, models like ChatGPT-4o and Qwen demonstrated social proof biases even when presented with non-indicative data, and displayed loss aversion when it was explicitly framed as a bias. Notably, the study reveals that academic papers documenting these cognitive biases can act as self-fulfilling prophecies, directly shaping the behavioral biases of LLMs exposed to them.


Document Metadata

Field Details
arXiv Identifier arXiv:2608.12339 [cs.CL]
Title Mimicry without understanding: the origins of decision bias in large language models
Authors Eldad Yechiam, Adi Tarabeih
Submitted On June 3, 2026
Primary Subject Computation and Language (cs.CL)
Secondary Subjects Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
DOI 10.48550/arXiv.2608.12339
Comments 33 pages, 3 figures, 2 boxes
Field Details
arXiv Identifier arXiv:2608.12339 [cs.CL]
Title Mimicry without understanding: the origins of decision bias in large language models
Authors Eldad Yechiam, Adi Tarabeih
Submitted On June 3, 2026
Primary Subject Computation and Language (cs.CL)
Secondary Subjects Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
DOI 10.48550/arXiv.2608.12339
Comments 33 pages, 3 figures, 2 boxes

Abstract

大语言模型(LLM)被发现容易受到各种社会、情感和认知偏差的影响。我们研究了产生此类偏差的两种机制,即使当人类偏好(在训练数据中)并不存在偏差,或者当它们被正确归类为带有偏差时:

  1. 对基于人类行为的偏好的错误模仿: 这涉及 LLM 对人类偏好的推断,即使这些行为在逻辑上与偏好无关。
  2. 对显式偏差人类行为的模仿: 这涉及对明确具有偏差的人类行为的直接模仿。

在聚焦于经济偏差的四项研究中,我们发现 ChatGPT-4o 和 Qwen 在被提示人类行为报告(这些报告明显无法指示个体的真实偏好)时,依然表现出社会认同偏差。当损失厌恶被明确描述为一种偏差时,LLM 也表现出了损失厌恶。事实上,当被输入详细的科学报告时,科学报告中偏差的程度(即损失厌恶程度)能够预测 LLM 自身随后的偏差。因此,关于偏差的科学论文可能会成为自证预言,至少在 LLM 的响应方面是如此。本研究不仅阐明了 LLM 的偏差,还揭示了其背后的底层成分加工过程。

Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training data) are not biased or when they are correctly categorized as being biased.

  1. Faulty mimicry of preferences based on human behavior: This involves LLMs inferring human preferences even when behaviors are logically unrelated to preferences.
  2. Mimicry of explicitly biased human behaviors: This involves the direct mimicry of explicitly biased human actions.

In four studies focusing on economic biases, we find that ChatGPT-4o and Qwen exhibited social proof biases even when prompted with reports of human behaviors that were clearly non-indicative of individuals' actual preferences. LLMs also displayed loss aversion when it was explicitly described as a bias. Indeed, when prompted with detailed scientific reports, the extent of the bias (i.e., loss aversion) in the scientific report predicted LLMs' own subsequent bias. Scientific papers of biases can thus become self-fulfilling prophecies, at least when it comes to LLMs' responses. The current study goes beyond fleshing out LLM biases and sheds light on the underlying component processes.


Additional Resources & Navigation