跳转至

部署决策可靠性:用于规划长周期智能体评估的泛化理论框架

文章背景与核心概要

企业实践者往往将智能体(Agent)排行榜误读为衡量总体智能体能力的权威排名。通过对三个开源智能体追踪基准(TheAgentCompany\(\tau^2\)-bench 和 AppWorld)的分析,本文证明了在所有数据集和检查类型中,单一智能体主效应占总方差的比例不到 3%。相反,“智能体与任务的交互作用”占了方差的 7% 到 23%,这表明排行榜实际上衡量的是任务专业化程度,而非通用能力。

为了解决这一问题,作者引入了一个四面(Four-facet)泛化理论(Generalizability Theory)方差分解框架,并采用了三种不同的估计器(Henderson 第一方法、基于 lme4 的 REML 以及贝叶斯二项式 GLMM),三者表现出高度的一致性。该研究提出了部署决策可靠性(Deployment Decision Reliability, DDR)——这是一个实用的报告框架,能够将复杂的方差分量表转化为五个可辩护的、面向企业采购和部署的决策。


Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

部署决策可靠性:用于规划长周期智能体评估的泛化理论框架

license icon

Summary

总结

Enterprise practitioners often misinterpret agent leaderboards as definitive rankings of overall agent capability. Analyzing three open agent-trace benchmarks (TheAgentCompany, \(\tau^2\)-bench, and AppWorld), this paper demonstrates that individual agent main effects account for less than 3% of the total variance across all datasets and check types. Instead, agent-by-task interactions account for 7% to 23% of the variance, revealing that leaderboards measure task specialization rather than general capability.

To address this, the author introduces a four-facet Generalizability Theory variance decomposition framework using three distinct estimators (Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM) that show high agreement. The study introduces Deployment Decision Reliability (DDR)—a practical reporting framework that translates complex variance component tables into five defensible, enterprise-ready procurement and deployment decisions.

企业从业者经常将智能体排行榜误解为总体智能体能力的决定性排名。通过分析三个开源智能体追踪基准(TheAgentCompany\(\tau^2\)-bench 和 AppWorld),本文证明了在所有数据集和检查类型中,个体智能体的主效应占总方差的比例不到 3%。相反,智能体与任务的交互作用占了方差的 7% 到 23%,这表明排行榜衡量的是任务专业化而非通用能力。

为了解决这一问题,作者引入了一个四面泛化理论方差分解框架,并使用三个不同的估计器(Henderson 第一方法、通过 lme4 的 REML 以及贝叶斯二项式 GLMM),这些估计器表现出高度的一致性。该研究引入了部署决策可靠性(Deployment Decision Reliability, DDR)——这是一个实用的报告框架,能将复杂的方差分量表转化为五个可辩护的、适合企业采购和部署的决策。


Article Metadata

文章元数据


Abstract

摘要

Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, \(\tau^2\)-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7-23%. Leaderboards rank specialization, not capability. We arrive at this through a four-facet Generalizability Theory variance decomposition, fit with three estimators (Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM) that agree to three decimal places.

Four further findings sharpen what the leaderboard is hiding: 1. Aggregate Reliability Collapse: Reliability collapses on the hardest task quartile (\(E\rho^2\) on \(\tau^2\) action_checks falls from 0.752 to 0.000). 2. Negative Training Correlation: Training-cell reliability negatively correlates with held-out reliability (\(r = -0.90\) on \(\tau^2\)), meaning the designs that look most reliable replicate worst. 3. Population Diagnostics Transfer: Population-level diagnostics transfer across enterprise benchmarks (capability-gap ratio stable at 0.35–0.40) while per-family agent rankings invert. 4. Failure Taxonomy Generalization: On the MAST failure taxonomy, trace-level mode profiles are idiosyncratic (\(\text{MAE} = 0.261\)) while cell-level profiles generalise (\(\text{MAE} = 0.056\), \(r = 0.83\)).

We package these into Deployment Decision Reliability (DDR), a one-page reporting discipline that turns the variance-component table into five decisions an enterprise buyer can defend. All code, data loaders, and fit artifacts are released under an open-source license.

企业从业者在阅读智能体排行榜时,往往将其视为智能体能力的排名。我们通过三个开源智能体追踪基准(TheAgentCompany、\(\tau^2\)-bench 和 AppWorld)发现,在每个数据集和检查类型中,智能体主效应占总方差的比例不到 3%,而智能体与任务的交互作用占 7-23%。排行榜对专业化进行了排名,而非能力。我们通过四面泛化理论方差分解得出了这一结论,并使用三个估计器(Henderson 第一方法、通过 lme4 的 REML 以及贝叶斯二项式 GLMM)进行拟合,它们在小数点后三位保持一致。

以下四个进一步的发现揭示了排行榜隐藏的内容: 1. 总体可靠性崩溃: 可靠性在最难的任务四分位数上崩溃(\(\tau^2\)action_checks 上的 \(E\rho^2\) 从 0.752 跌至 0.000)。 2. 负向训练相关性: 训练单元可靠性与留出(held-out)可靠性呈负相关(在 \(\tau^2\)\(r = -0.90\)),这意味着看起来最可靠的设计在重复测试时表现最差。 3. 群体诊断迁移: 群体级诊断可在企业基准之间迁移(能力差距比率保持在 0.35–0.40 稳定),而各家族的智能体排名则发生倒置。 4. 失败分类泛化: 在 MAST 失败分类法中,追踪级模式配置文件具有特异性(\(\text{MAE} = 0.261\)),而单元级配置文件则具有通用性(\(\text{MAE} = 0.056\), \(r = 0.83\))。

我们将这些成果打包为部署决策可靠性(Deployment Decision Reliability, DDR),这是一种单页报告规范,可将方差分量表转化为企业买家能够辩护的五项决策。所有代码、数据加载器和拟合工件均已在开源许可证下发布。


Access and Resources

访问与资源