arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11323cs.AIcs.LG

部署决策可靠性:用于评估长 horizon 智能体规模的概化理论框架

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

Vasundra Srinivasan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究采用概化理论框架分析三个智能体基准的排行榜,发现其排名反映专业化而非能力,提出 DDR 报告规范并发布相关资源,为企业评估长 horizon 智能体提供可靠方法。

中文摘要 AI 辅助

企业从业者将智能体排行榜视为对智能体能力的排名。我们在三个开放智能体轨迹基准(TheAgentCompany、τ²-bench 和 AppWorld)中发现,在每个数据集和检查类型下,智能体主效应对总方差的贡献不足 3%,而智能体与任务的交互效应贡献为 7-23%。排行榜排名的是专业化程度,而非能力。我们通过四维度概化理论方差分解得出此结论,该分解采用三种估计量(Henderson Method-I、通过 lme4 的 REML 以及贝叶斯二项 GLMM),三者结果在小数点后三位一致。另有四项发现揭示了排行榜隐藏的信息:其一,综合可靠性在最难任务四分位数上崩溃:τ² 的 action_checks 的 Eρ² 从 0.752 降至 0.000;其二,训练单元可靠性与保留可靠性呈负相关(τ² 上 r = -0.90),即看似最可靠的设计复制效果最差;其三,总体水平诊断可跨企业基准迁移(能力差距比稳定在 0.35-0.40),但按家族划分的智能体排名会反转;其四,在 MAST 故障分类法上,轨迹级模式特征具有特殊性(MAE = 0.261),而单元级特征可泛化(MAE = 0.056,r = 0.83)。我们将这些结果整合为部署决策可靠性(DDR),这是一种一页报告规范,可将方差成分表转化为企业买家可辩护的五项决策。所有代码、数据加载器和拟合工件均以开源许可发布。

英文摘要

Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $τ^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7-23%. Leaderboards rank specialization, not capability. We arrive at this through a four-facet Generalizability Theory variance decomposition, fit with three estimators (Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM) that agree to three decimal places. Four further findings sharpen what the leaderboard is hiding. First, aggregate reliability collapses on the hardest task quartile: $Eρ^2$ on $τ^2$ action_checks falls from 0.752 to 0.000. Second, training-cell reliability negatively correlates with held-out reliability ($r = -0.90$ on $τ^2$), meaning the designs that look most reliable replicate worst. Third, population-level diagnostics transfer across enterprise benchmarks (capability-gap ratio stable at 0.35-0.40) but per-family agent rankings invert. Fourth, on the MAST failure taxonomy, trace-level mode profiles are idiosyncratic (MAE = 0.261) while cell-level profiles generalise (MAE = 0.056, $r = 0.83$). We package these into Deployment Decision Reliability (DDR), a one-page reporting discipline that turns the variance-component table into five decisions an enterprise buyer can defend. All code, data loaders, and fit artifacts are released under an open-source license.

发表机构

  • Stanford School of Engineering(斯坦福工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑