跳转至

人类心理测量问卷错误表征了大语言模型的行为

文章背景与核心概要

近年来,研究人员越来越多地借用人类心理学中的心理测量问卷(如价值观问卷和性格量表)来评估大语言模型(LLM)的“个性”与“价值观”。然而,这些标准化测试是否能够真正反映模型在日常用户交互中的实际行为,一直缺乏系统的检验。本文作者通过对八个开源大语言模型进行深入研究,对比了问卷自我报告结果与生态有效查询下的实际生成概率,揭示了人类心理测量问卷在刻画LLM行为方面的局限性。

研究发现,问卷中的显式词汇线索使得模型能够识别出目标构建并输出符合对齐要求的社会赞许性回答,但这种“稳定的人格特征”在真实的生成任务中会完全消失。此外,尽管人口统计学角色提示词能让模型在人类问卷上表现出符合人类模式的偏移,但在实际生成场景中却毫无效果。这表明人类心理测量分数高估了LLM的心理保真度,作者因此倡导采用基于生态有效题目的“生成概率分析法”(generation-probability profiling)作为更准确的行为评估指标。


摘要 (Abstract)

我们探讨了人类心理测量问卷是否能作为可靠的工具,来表征和预测大语言模型(LLM)在日常用户交互中的行为。通过分析八个开源大语言模型,我们对比了通过两种不同方法得出的其价值观和个性画像:一是基于既定问卷(PVQ-40/21 和 BFI-44/10)的李克特量表自我报告,二是针对日常用户查询中蕴含价值取向回答的生成概率。

We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open-source LLMs by comparing their value and personality profiles derived from two different methods: Likert self-reports on established questionnaires (PVQ-40/21 and BFI-44/10) and generation probabilities over value-laden responses to everyday user queries.

这两种画像存在显著的分歧。常被引作LLM具有稳定性格倾向证据的构念内题目一致性,在生成概率中荡然无存。我们发现,既定的问卷题目包含显式的词汇线索,使模型能够识别目标构念并以符合对齐、具有社会赞许性的方式进行响应,而现实中的用户查询包含的明显线索则少得多。

The two profiles diverge substantially. Within-construct item consistency, often cited as evidence of stable LLM dispositions, disappears in generation probabilities. We find that established questionnaire items contain explicit lexical cues that allow models to recognize the target construct and respond in alignment-consistent, socially desirable ways, whereas realistic user queries contain far less recognizable cues.

此外,人口统计学角色提示词会以符合人类真实模式的方式改变模型对人类问卷的响应,但在生成场景中并没有出现这种偏移,这凸显出人类心理测量高估了LLM在扮演人口统计学角色时忠实重现预期心理特征的能力。总体而言,我们的研究表明,不应将问卷得分单独视为LLM在现实用户交互中响应倾向的证据,并支持将利用生态有效题目进行的生成概率分析作为一种互补的行为测量手段。

In addition, demographic persona prompts shift models' responses to human questionnaires in ways consistent with real human patterns, but no such shifts appear in the generation setting, highlighting that human questionnaires overestimate LLMs' ability to faithfully reproduce expected psychological traits when role-playing demographic personas. Overall, our study indicates that questionnaire scores alone should not be treated as evidence of LLMs' response tendencies in realistic user interactions, and supports generation-probability profiling with ecologically valid items as a complementary behavioral measure.


元数据 (Metadata)

  • arXiv ID: arXiv:2509.10078 [cs.CL]
  • DOI: 10.48550/arXiv.2509.10078
  • 学科分类 (Subjects): 计算与语言 (cs.CL); 人工智能 (cs.AI)
  • 会议信息 (Conference): 已被 EMNLP 2026(主会)接收
  • 提交历史 (Submission History):
  • 2025年9月12日提交 (v1)
  • 2026年9月2日最后修订 (v5)
  • arXiv ID: arXiv:2509.10078 [cs.CL]
  • DOI: 10.48550/arXiv.2509.10078
  • Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
  • Conference: Accepted at EMNLP 2026 (Main)
  • Submission History:
  • Submitted on 12 Sep 2025 (v1)
  • Last revised on 2 Sep 2026 (v5)

作者 (Authors)

  • Woojung Song
  • Dongmin Choi
  • Yoonah Park
  • Jongwook Han
  • Eun-Ju Lee
  • Yohan Jo
  • Woojung Song
  • Dongmin Choi
  • Yoonah Park
  • Jongwook Han
  • Eun-Ju Lee
  • Yohan Jo