跳转至

智能体物理学:统计力学预测人工智能群体的集体行为

文章背景与核心概要

随着人工智能从孤立的模型演变为相互作用的多智能体系统,理解其集体动力学对于确保系统有效对齐至关重要。本文通过对超过 10,000 个语言模型智能体社区进行研究,观察了它们在客观数学问题和主观政治议题中交换信息并修正观点的过程。

研究发现,AI 的集体行为表现出与物理复杂系统类似的特征,自然地归纳为三种动力学机制:冷漠(indifference)极化(polarization)共识(consensus)。作者引入了一种统计力学框架,将智能体建模为随机最小化社会压力的主体。该预测模型不仅在性能上优于标准基准,还能泛化至未见的社区图谱,并揭示了关键的力学洞察,例如社区运行在临界社会温度之下、吸引力纽带强于排斥力纽带,以及追求真理的智能体具有最强的集体拉力。


文档元数据

  • arXiv 标识符: arXiv:2608.16578 [cs.AI]
  • 提交日期: 2026年8月17日
  • 主要学科: 人工智能 (cs.AI)
  • 次要学科: 多智能体系统 (cs.MA);社会与信息网络 (cs.SI)
  • 篇幅: 51页,20张图表,9个表格
  • 许可协议: 知识共享署名 4.0 国际许可协议 license icon

作者

  • Batu El
  • Jinhee Paeng
  • Fatih Dinc
  • Shiye Su
  • Mete Erdogan
  • Aneesh Pappu
  • Haotian Ye
  • Wanjia Zhao
  • Surya Ganguli
  • James Zou

Authors

  • Batu El
  • Jinhee Paeng
  • Fatih Dinc
  • Shiye Su
  • Mete Erdogan
  • Aneesh Pappu
  • Haotian Ye
  • Wanjia Zhao
  • Surya Ganguli
  • James Zou

摘要

人工智能智能体越来越多地作为交互系统的一部分运行,而非孤立存在。当智能体交换信息并共同做出决策时,它们的互动可以改善集体推理,但也可能产生从众、极化或放大共享偏见。因此,理解并预测这些集体动力学对于设计有效且对齐的多智能体系统非常重要。

在此,我们研究了超过 10,000 个语言模型智能体社区,它们在客观数学问题和主观政治陈述中反复交换信息并修正观点。尽管可能的行为存在巨大差异,但个体和群体的动力学可以用三种特征机制来表示:冷漠极化共识

Abstract

AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems.

Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus.

AI 智能体起初表现冷漠,并在互动中建立信念: * 在客观问题上,沟通提高了集体准确性。 * 在主观问题上,它往往使群体观点向政治光谱的右侧偏移。

AI agents start indifferent and build conviction as they interact: * On objective questions, communication improves collective accuracy. * On subjective questions, it often drifts group opinions toward the right in the political spectrum.

我们用一种统计力学形式体系解释了这些观察结果,在该体系中,智能体随机地倾向于较低的社会压力。仅给定初始观点,我们的模型能够: 1. 预测个体轨迹并优于所有标准基准。 2. 泛化到未见的社区图谱。 3. 再现观察到的群体原型分布。

We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure. Given only initial opinions, our model: 1. Predicts individual trajectories and outperforms all standard baselines. 2. Generalizes to unseen community graphs. 3. Reproduces the observed group archetype distributions.

我们拟合的模型参数揭示了上述关键观察背后的力学机制: * i) 社区运行在临界社会温度之下,这解释了信念的建立。 * ii) 吸引力纽带强于排斥力纽带,这有利于达成共识。 * iii) 持有正确答案的智能体施加了最强的拉力,这推动了对真理的追求。

总体而言,我们的结果表明,AI 智能体的集体行为与其它复杂系统一样,遵循紧凑且具有预测性的动力学定律。

Our fitted model parameters reveal the mechanics underlying our key observations: * i) Communities operate below the critical social temperature, which explains conviction buildup. * ii) Attractive ties outweigh repulsive ones, which favors consensus. * iii) Agents holding the correct answer exert the strongest pull, which drives truth-seeking.

Overall, our results demonstrate that the collective behavior of AI agents, like that of other complex systems, follows compact and predictive dynamical laws.


获取与资源

Access & Resources