当大模型智能体进行博弈:供应链中的私有信息与动态讨价还价
文章背景与核心概要
随着大语言模型(LLM)智能体从决策支持工具演变为自主采购实体,企业正面临着关于其效能的关键问题:被委托的谈判代表是否能创造价值、以可预测的方式分配价值,并避开亏损合同?本文研究了一个经典的供应链讨价还价问题:拥有私有需求信息的买方与不知情的主体(卖方)就数量-支付合同进行谈判。
通过在 9,840 场大模型对大模型的谈判中,将来自 OpenAI、谷歌和阿里巴巴的九大主流大模型与经过验证的“完美贝叶斯均衡”(Perfect Bayesian Equilibrium)进行基准对比,作者从折扣效率(discounted efficiency)、分配特征(distributional profile)和运营可靠性(operational reliability)三个主要维度评估了人工智能的表现。研究发现,虽然智能体达成了极高的协议成功率,但其冗长的谈判轮次会侵蚀大量剩余价值,且模型供应商的身份对价值分配具有显著影响。
执行摘要 / Executive Summary
As Large Language Model (LLM) agents transition from decision support tools to autonomous procurement entities, firms face critical questions regarding their utility: Do delegated negotiators create value, divide it predictably, and steer clear of money-losing contracts?
随着大语言模型(LLM)智能体从决策支持工具演变为自主采购实体,企业正面临着关于其效能的关键问题:被委托的谈判代表是否能创造价值、以可预测的方式分配价值,并避开亏损合同?
This study examines a canonical supply chain bargaining problem where a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller. Benchmarking nine major LLMs (from OpenAI, Google, and Alibaba) against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations, the authors evaluate AI performance across three primary dimensions: discounted efficiency, distributional profile, and operational reliability.
本研究探讨了一个经典的供应链讨价还价问题:拥有私有需求信息的买方与不知情的卖方就数量-支付合同进行谈判。通过在 9,840 场大模型对大模型的谈判中,将九大主流大模型(来自 OpenAI、谷歌和阿里巴巴)与经过验证的完美贝叶斯均衡(Perfect Bayesian Equilibrium)进行基准对比,作者从三个主要维度评估了人工智能的表现:折扣效率、分配特征和运营可靠性。
核心发现 / Key Findings
1. 能力决定价值创造与可靠性 / 1. Capability Governs Value Creation and Reliability
- Efficiency & Delay: Agents successfully agree in 98.9% of negotiations, capturing 95.4% of the first-best surplus (undiscounted). However, they average 2.98 rounds of negotiation compared to the benchmark's 1.25 rounds—a delay that erodes 21–34% of the surplus.
- 效率与延迟: 智能体在 98.9% 的谈判中成功达成协议,捕获了 95.4% 的一阶最优剩余(未折现)。然而,它们的平均谈判轮次为 2.98 轮,而基准模型仅为 1.25 轮——这种延迟侵蚀了 21% 至 34% 的剩余价值。
- Rationality & Guardrails: Baseline models accept individually irrational (money-losing) contracts in 19.2% of cases, whereas mid-tier and flagship models drop this rate to 0.0–0.6%. Consequently, automated profit verification serves as a necessary binding guardrail for lower-tier models.
- 理性与防护栏: 基线模型在 19.2% 的情况下会接受单方面不理性(亏损)的合同,而中端和旗舰模型的这一比例降至 0.0% 至 0.6%。因此,自动化利润验证对于低端模型而言是一道必需的约束性防护栏。
2. 剩余分配具有关系性(供应商身份至关重要) / 2. Surplus Capture is Relational (Provider Identity Matters)
- Provider Bias: A model's provider identity is a stronger predictor of surplus capture than its general capability rank. For instance, self-play buyer shares average 40% for OpenAI, 50% for Google, and 70% for Alibaba's Qwen—an ordering that persists even with restricted communication and no discounting.
- 供应商偏见: 模型供应商的身份比其通用能力排名更能预测剩余价值的捕获情况。例如,在自我对弈(self-play)中,买方份额在 OpenAI 中平均为 40%,谷歌为 50%,而阿里巴巴的通义千问(Qwen)则高达 70%——即使在限制通信且无折现的情况下,这种排序依然存在。
- Distributional Shifts: Reversing roles (who is selling) shifts the financial division by 7 to 18 percentage points. Interestingly, the highly capable Qwen flagship model functions as the weakest cross-family seller, proving that vendor choice is a fundamental distributional decision.
- 分配转移: 角色反转(谁在出售)会使财务分配发生 7 到 18 个百分点的偏移。有趣的是,能力极强的 Qwen 旗舰模型在跨家族销售中表现为最弱的卖方,这证明了供应商的选择是一个根本性的分配决策。
3. 提示词是战略杠杆 / 3. The Prompt is a Strategic Lever
- Strategic Patience: Delegation cleanly separates the principal's economic patience from the agent's prompted strategic patience.
- 战略耐心: 委托机制清晰地将委托人的经济耐心与智能体被提示的战略耐心区分开来。
- Variance Impact: This free deployment choice acts as the single strongest driver of surplus division, accounting for 90% of explained variance.
- 方差影响: 这种免费的部署选择成为剩余分配的最强驱动力,占到了解释方差的 90%。
结论 / Conclusion
This research establishes an equilibrium-referenced audit framework for AI agents in commercial negotiations. By mapping out discounted efficiencies, distributional profiles, and operational reliabilities, firms can better govern autonomous multi-agent procurement systems and optimize deployment parameters.
本研究为商业谈判中的人工智能智能体建立了一个以均衡为参考的审计框架。通过绘制折扣效率、分配特征和运营可靠性的蓝图,企业可以更好地治理自主多智能体采购系统并优化部署参数。