大语言模型角色双重性:聚合倾向与框架依赖几何学
文章背景与核心概要
传统的心理测量学问卷评估大语言模型(LLM)角色时,通常严重依赖聚合得分(例如大五人格特质评估),这往往会丢弃至关重要的实例内部相关性结构。本文旨在探讨这种几何结构究竟是一种内在特质,还是仅仅取决于提问框架。
作者通过在对称正定(SPD)流形上分析来自IPIP-50反应的实例内部相关矩阵——利用GPT-4o在随机化问题顺序下模拟美国和美籍华人角色——证明了LLM角色的表达由两个截然不同且可分离的组件驱动: 1. 聚合特征(框架鲁棒型): 标准的聚合得分(如大五人格特质)在随机化下略有下降(下降21%),但在不同的框架背景下保持稳健。 2. 几何特征(框架依赖型): SPD流形配置在框架不对齐时显著崩溃(下降42%),但在共享框架下表现出惊人的恢复能力(高达84%),轻松超越了聚合特征(76%)。
这种“崩溃-恢复”的动态变化揭示了角色几何学并非内在的静态特质,而是一种框架依赖的协调模式,它编码了传统基础聚合完全无法捕捉的有价值信息。这些发现呼吁人工智能领域向具备框架感知能力(frame-aware)的评估框架转变。
📌 摘要与核心内容 (Summary)
传统 psychometric evaluations of Large Language Model (LLM) personas rely heavily on aggregate scores (such as Big Five trait assessments), which often discard the critical within-instance correlation structure. This paper investigates whether this geometric structure is an intrinsic trait or merely frame-dependent.
By analyzing within-instance correlation matrices from IPIP-50 responses across symmetric positive definite (SPD) manifolds—using GPT-4o to simulate American and Chinese-American personas under randomized question orderings—the author demonstrates that LLM persona expression is driven by two distinct, dissociable components: 1. Aggregated Features (Frame-Robust): Standard aggregate scores (e.g., Big Five traits) degrade slightly under randomization (a 21% drop) but remain robust across different framing contexts. 2. Geometric Features (Frame-Dependent): SPD manifold configurations collapse significantly under frame misalignment (a 42% drop), yet recover remarkably well (up to 84%) under shared frames, easily outperforming aggregated features (76%).
This collapse-recovery dynamic reveals that persona geometry is not an intrinsic, static trait, but rather a frame-dependent coordination pattern that encodes valuable information entirely invisible to basic aggregation. These findings advocate for a shift toward frame-aware evaluation frameworks in artificial intelligence.
传统心理测量问卷对大语言模型(LLM)角色的评估通常依赖聚合得分,从而丢弃了实例内部的相关性结构。本文测试了几何结构是内在特质还是依赖于提问框架。通过构建IPIP-50反应的实例内部相关矩阵,并在GPT-4o模拟美国和美籍华人角色时操纵问题顺序,在对称正定(SPD)流形上分析了几何结构。我们发现角色表达包含两个可分离的组件:聚合特征(大五人格得分)在随机化下会退化(下降21%),但具有框架鲁棒性;几何特征(SPD流形)在框架不对齐时崩溃(下降42%),但在共享框架下能大幅恢复(至84%),超越了聚合特征(76%)。这种崩溃-恢复模式表明,角色几何学不是内在的,而是一种框架依赖的协调模式,它编码了聚合所不可见的隐式信息。我们的研究结果确立了LLM角色的双重性质框架(框架依赖几何学与框架鲁棒聚合),凸显了采用框架感知评估的必要性,并对静态特质观念提出了挑战。
🔗 链接与资源 (Links & Resources)
- Full-Text Access: View PDF | HTML (Experimental) | TeX Source
- DOI: 10.48550/arXiv.2607.02368
- License: Creative Commons Attribution-NonCommercial-NoDerivatives 4.0
- External Citations: Google Scholar | Semantic Scholar | NASA ADS
- 全文访问: 查看 PDF | HTML(实验性) | TeX 源码
- DOI: 10.48550/arXiv.2607.02368
- 许可证: 知识共享 署名-非商业性使用-禁止演绎 4.0
- 外部引用: Google Scholar | Semantic Scholar | NASA ADS
📜 摘要 (Abstract)
Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation structure. We test whether this geometric structure is intrinsic or frame-dependent. Constructing within-instance correlation matrices from IPIP-50 responses, we analyze geometry on SPD manifolds under manipulated question orderings in GPT-4o simulating American and Chinese-American personas. We find that persona expression comprises two dissociable components: aggregated features (Big Five scores) degrade under randomization (21% drop) but are frame-robust; geometric features (SPD manifold) collapse under frame misalignment (42% drop) but recover substantially (to 84%) under shared frames, surpassing aggregated features (76%). This collapse-recovery pattern reveals that persona geometry is not intrinsic but a frame-dependent coordination pattern encoding information invisible to aggregation. Our findings establish a dual-nature framework for LLM personas, frame-dependent geometry versus frame-robust aggregates, necessitating frame-aware evaluation and challenging static trait conceptions.
通过心理测量问卷评估LLM角色通常依赖于聚合得分,从而丢弃了实例内部的相关结构。我们测试了这种几何结构是内在的还是框架依赖的。通过构建IPIP-50反应的实例内部相关矩阵,我们在GPT-4o模拟美国和美籍华人角色时,在操纵问题顺序的情况下分析了SPD流形上的几何结构。我们发现角色表达包含两个可分离的组件:聚合特征(大五人格得分)在随机化下会退化(下降21%),但具有框架鲁棒性;几何特征(SPD流形)在框架不对齐时崩溃(下降42%),但在共享框架下显著恢复(至84%),超过了聚合特征(76%)。这种崩溃-恢复模式表明,角色几何学不是内在的,而是一种框架依赖的协调模式,它编码了聚合所不可见的冗余或隐式信息。我们的研究结果为LLM角色确立了一个双重性质框架(框架依赖几何学与框架鲁棒聚合),这使得框架感知的评估变得必不可少,并对静态特质的概念构成了挑战。
🗂️ 提交历史 (Submission History)
- [v1] Thu, 2 Jul 2026, 16:11:44 UTC (646 KB)
- [v2] Fri, 21 Aug 2026, 14:37:51 UTC (147 KB) — Current Version
- [v1] 2026年7月2日 星期四,16:11:44 UTC (646 KB)
- [v2] 2026年8月21日 星期五,14:37:51 UTC (147 KB) — 当前版本