跳转至

随机鹦鹉还是和谐共鸣?测试五大主流大模型利用合成数据复刻人类调查的能力

文章背景与核心概要

本研究探讨了人工智能生成的合成数据在多大程度上能够成功复刻组织研究中人类参与者的真实反馈。作者选取了硅谷420名程序员和开发人员的真实调查数据,并将其与由五种主流大语言模型(ChatGPT Thinking 5 Pro、Claude Sonnet 4.5 Pro + CoWork 1.123、Gemini Advanced 2.5 Pro、Incredible 1.0 以及 DeepSeek 3.2)生成的合成调查数据进行了对比分析。

研究揭示了一个关键悖论:尽管所有受测模型都能生成看起来合理且彼此高度一致的数据,但它们均未能捕捉到原始人类调查中那些反直觉、细致入微的洞察。模型表现得更像是“随机鹦鹉”,反映的是传统认知而非真实的人类社会信念。因此,作者得出结论:合成调查数据不应取代严谨的实地调研,而应作为一种辅助工具,用于在开展人类研究前后识别社会假设和预期。


📌 总结

This research paper investigates whether AI-generated synthetic data can successfully replicate the real-world responses of human participants in organizational research. The authors compared a human-respondent survey of 420 Silicon Valley coders and developers against synthetic survey data generated by five leading Large Language Models (LLMs): ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro + CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2.

The study reveals a critical paradox: while all tested LLMs produced plausible data that harmonized well with one another, none were able to capture the counterintuitive, nuanced insights that made the original human survey truly valuable. Instead, the models acted as "stochastic parrots," reflecting conventional wisdom rather than real human social beliefs. Consequently, the authors conclude that synthetic survey data should not replace rigorous fieldwork, but rather serve as a supplementary tool for identifying societal assumptions and expectations before or after conducting human research.


📑 元数据与参考链接


📝 摘要

人工智能生成的合成研究数据在多大程度上能够复刻人类参与者的反馈?新兴文献已开始关注这一问题,这对组织研究实践具有深远影响。本文对比了硅谷420名程序员和开发人员的人类受访者调查数据,以及由五种主流生成式人工智能大模型(ChatGPT Thinking 5 Pro、Claude Sonnet 4.5 Pro + Claude CoWork 1.123、Gemini Advanced 2.5 Pro、Incredible 1.0 和 DeepSeek 3.2)生成的旨在模拟真实受访者的合成调查数据。

How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article presents a comparison between a human-respondent survey of 420 Silicon Valley coders and developers and synthetic survey data designed to simulate real survey takers generated by five leading Generative AI Large Language Models: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2.

我们的研究结果显示,虽然人工智能代理生成的结论在技术上看似合理,且在可复刻性和和谐度上超出了预期,但没有一个模型能够捕捉到使人类调查具有价值的反直觉洞察。此外,所有模型的偏差趋于一致,使得真实数据反而成为了离群值。

Our findings reveal that while AI agents produced technically plausible results that lean more towards replicability and harmonization than assumed, none were able to capture the counterintuitive insights that made the human survey valuable. Moreover, deviations grouped together for all models, leaving the real data as the outlier.

我们的核心发现是,尽管领先的大语言模型正越来越多地被用于扩展、复刻和替代研究中的人类调查反馈,但这些进展仅表明它们在和谐地复述传统认知方面能力增强,而非揭示了新的发现。如果未来研究中使用合成受访者,我们需要更具可复刻性的验证协议和报告标准,以明确合成调查数据在何时何地可以被负责任地使用,而本文填补了这一空白。我们的结果表明,合成调查反馈无法有效地模拟组织内真实的人类社会信念,特别是在缺乏既往记录证据的环境中。我们得出结论:基于合成调查的研究不应被视为严谨调查方法的替代品,而应作为一种日益可靠的实地调研前或调研后工具,用于识别社会假设、传统认知以及对研究群体的其他预期。

Our key finding is that while leading LLMs are increasingly being used to scale, replicate and replace human survey responses in research, these advances only show an increased capacity to parrot conventional wisdom in harmony with each other rather than revealing novel findings. If synthetic respondents are used in future research, we need more replicable validation protocols and reporting standards for when and where synthetic survey data can be used responsibly, a gap that this paper fills. Our results suggest that synthetic survey responses cannot meaningfully model real human social beliefs within organizations, particularly in contexts lacking previously documented evidence. We conclude that synthetic survey-based research should be cast not as a substitute for rigorous survey methods, but as an increasingly reliable pre- or post-fieldwork instrument for identifying societal assumptions, conventional wisdoms, and other expectations about research populations.