文章背景与核心概要
在多智能体系统中,当具有社会价值的努力成本高昂、难以观测且主要使其他参与者受益时,合作往往会瓦解。本文借鉴霍姆斯特姆(Holmström)关于团队道德风险的经典经济学模型,引入了对话道德风险博弈(Dialogue Moral Hazard Game)。这一基于理论且可控的实验框架将隐蔽行动结构转化为适合语言智能体的文本环境。
通过对11个开源模型和3个前沿API模型的各项评估,该研究分析了查询率、信息传递、局部奖励保留、不安全选择以及整体团队成功率等指标。研究结果表明,标准优化技术(如SFT、RLOO以及通过GEPA进行的提示词优化)可以在不一定恢复底层合作机制的情况下提升总体奖励——这凸显了评估机制层面行为而非仅仅关注团队成功率的极端重要性。
Moral Hazard in Multi-Agent Language Models
Authors: Dane Malenfant
Subjects: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
arXiv: 2607.23982 [cs.MA]
DOI: 10.48550/arXiv.2607.23982
Publication History: Submitted July 27, 2026; Last revised August 13, 2026 (v4)
📌 Summary
Cooperation in multi-agent systems frequently breaks down when socially valuable effort is costly, difficult to observe, and primarily benefits other participants. Drawing inspiration from Holmström’s classic economic model of moral hazard in teams, this paper introduces the Dialogue Moral Hazard Game. This theory-grounded, controlled experimental framework transforms hidden-action structures into a textual environment suitable for language agents.
Across various evaluations involving eleven open-weight models and three frontier API models, the study analyzes metrics such as query rates, information transfer, local-reward preservation, unsafe choices, and overall team success. The findings reveal that standard optimization techniques (such as SFT, RLOO, and prompt optimization via GEPA) can boost aggregate rewards without necessarily restoring the underlying cooperative mechanism—highlighting the critical need to evaluate mechanism-level behavior rather than team success alone.
🔬 Key Experimental Findings
- Frontier Model Variability: Different frontier policies display distinct incentive-balancing behaviors. For instance, Fable 5 dynamically adjusts its querying based on query cost and team reward, whereas GPT-5.6 Sol reaches behavioral ceilings in primary settings.
- Incentive Isolation: In a 3,015-decision incentive-isolation experiment, GPT-5.6 Sol closely tracked the Holmström-derived private-share boundary across nine distinct query costs with a remarkably low mean absolute error of
0.013.- Impact of Optimization: Supervised fine-tuning (SFT), REINFORCE Leave-One-Out (RLOO), and GEPA prompt optimization produced heterogeneous effects. While models like SmolLM3-3B and OLMo-7B demonstrated clear mechanism-consistent gains, GEPA occasionally increased team success while entirely bypassing or eliminating costly queries.
🔗 Links & Resources
- Full-Text Access: View PDF | HTML Version | TeX Source
- License: Creative Commons Attribution 4.0
- Citation Tools: NASA ADS | Google Scholar | Semantic Scholar
(Note: License icon preserved per source requirement:
)
)