文章背景与核心概要
随着大语言模型(LLM)向深度推理和复杂任务演进,推理时代币(Reasoning-Token)的分配成为了影响用户体验和平台收益的核心问题。服务商需要在设定单价和默认推理代币预算之间做出权衡,而用户则面临着接受默认配置、自行定制或直接弃用服务的选择。
本文运用博弈论框架,深入研究了LLM服务商与用户之间关于“推理代币”分配的经济动态。研究发现,用户的最优定制方案具有唯一的闭式解,而服务商的最优默认设置则遵循“三区制规则(three-regime rule)”,从而将复杂的市场均衡简化为一维的价格优化问题。实验表明,通过在多个数学和科学基准测试上对开源推理模型进行验证,该研究不仅证实了准确度-代币模型的有效性,还揭示了特定任务特征如何直接决定定价与服务设计。
保留、定制还是退出:LLM推理服务中的默认设计与代币定价 (Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services)
作者: Ahmet Bugra Gundogan, Yigit Turkmen, Melih Bastopcu
日期: 2026年8月13日
arXiv ID: 2608.13315
主要学科: 计算机科学与博弈论 (cs.GT)
摘要 (Summary)
本文研究了LLM服务商与用户之间关于“推理代币”分配的经济动态。服务商必须设定每个代币的价格以及默认的推理代币预算,而用户则决定是接受默认配置、定制分配方案,还是完全退出服务。
作者将这种互动建模为斯塔克伯格博弈(Stackelberg game),并得出以下结论: * 用户行为: 用户的最优定制方案存在唯一的闭式解(closed-form)。 * 服务商策略: 服务商的最优默认设置遵循三区制规则,且市场均衡可简化为一维的价格优化问题。 * 便利性因素: 只有当用户重视避免手动定制的便利性时,默认设置才会影响最终的推理代币分配;否则,市场自然会趋向于用户的最优分配。 * 实证验证: 在五个数学和科学基准测试上使用开源推理模型的实验,证实了准确度-代币模型的有效性,并阐明了特定任务特征如何决定定价与设计。
This paper investigates the economic dynamics between LLM service providers and users regarding "reasoning-token" allocation. Providers must set a per-token price and a default reasoning-token budget, while users decide whether to accept the default, customize their allocation, or exit the service entirely.
The authors model this interaction as a Stackelberg game, finding that: * User Behavior: There is a unique, closed-form optimal customization for users. * Provider Strategy: The provider’s optimal default follows a three-regime rule, and the equilibrium can be reduced to a one-dimensional price optimization. * Convenience Factor: Defaults only influence the final reasoning allocation when users value the convenience of avoiding manual customization. Otherwise, the market naturally gravitates toward the user's optimal allocation. * Empirical Validation: Experiments using open-weight reasoning models across five mathematics and science benchmarks confirm the validity of the accuracy-token model and illustrate how specific task characteristics dictate pricing and design.
核心研究发现 (Key Research Findings)
1. 博弈论框架
该研究将LLM服务视为一种战略互动,其中: * 更高的分配可以提高模型准确率,但同时也会增加成本和延迟。 * 服务商旨在优化定价和默认设置以实现效用最大化。 * 用户则在准确率与资源消耗之间进行权衡。
1. The Game-Theoretic Framework
The study treats the LLM service as a strategic interaction where: * Higher allocations improve model accuracy but increase costs and latency. * The Provider aims to optimize pricing and defaults to maximize utility. * The User balances the trade-off between accuracy and resource expenditure.
2. 均衡动态
研究人员证明了该服务模型中均衡的存在性。他们证明了对于任何给定的价格,可接受的默认值集合要么为空,要么形成一个紧区间。这使服务商能够策略性地设定默认值,从而引导用户采取特定行为,或适应用户对便利性的不同偏好。
2. Equilibrium Dynamics
The researchers proved the existence of an equilibrium in this service model. They demonstrated that for any given price, the set of acceptable defaults is either empty or forms a compact interval. This allows providers to strategically set defaults that either nudge users toward specific behaviors or accommodate varying user preferences for convenience.
3. 实际意义
本文为LLM服务商提供了一个行动指南,帮助其理解: * 如何设定默认推理预算以最大化用户留存率。 * 如何根据模型的特定“推理-准确率”曲线对代币进行有效定价。 * 用户“便利性成本”对整体服务架构的影响。
3. Practical Implications
The paper provides a roadmap for LLM providers to understand: * How to set default reasoning budgets to maximize retention. * How to price tokens effectively based on the specific "reasoning-to-accuracy" curve of their models. * The impact of user "convenience costs" on the overall service architecture.
访问与资源 (Access & Resources)
Access & Resources