风险感知的自适应评估:在有限预算下发现高影响失败
Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对交互式智能体评估预算有限且失败罕见的问题,提出风险感知的上下文汤普森采样策略,通过顺序分配试验预算,在最小预算下恢复86%的高影响失败,显著优于均匀分配。
AI中文摘要:
评估交互式智能体成本高昂。智能体行为具有随机性,因此可靠性必须通过重复试验来度量,但失败是罕见的,且其重要性差异巨大。标准基准测试均匀地分配预算:只读查找与不可逆的支付操作被同等频率地采样。我们转而将评估表述为一个顺序分配问题。给定固定的试验预算和一组失败行为未知的场景,应该运行哪些场景,以及再次运行哪些场景?我们提出了一种风险感知的上下文汤普森采样策略,该策略将执行前的场景上下文向量和固定的影响评分与评估期间观察到的失败结果相结合,并通过在70个τ-bench航空场景和824条记录试验上的离线重放进行测试。我们的主要结果出现在最小预算下:仅用50次试验(占语料库的6%),该策略恢复了预言机所能发现的影响加权失败的86%,而均匀分配仅为25%。在相同试验次数下,它发现的影响加权失败数量是均匀分配的3.5倍(215.4对62.2),每美元发现量是其5倍,并将浪费在从未失败场景上的预算从34%削减至2.8%。我们分析的其余部分展示并限定了这一结果:预算扫描显示,随着预算接近语料库规模,优势缩小;配对显著性检验表明,场景上下文主要在预算较小时有帮助,而基于后验的探索在中等预算时有帮助。因此,风险感知的自适应分配在评估预算最稀缺时帮助最大。
英文摘要:
Evaluating interactive agents is expensive. Agent behavior is stochastic, so reliability must be measured over repeated trials, but failures are rare and differ widely in how much they matter. Standard benchmarks spend this budget uniformly: a read-only lookup is sampled as often as an irreversible payment action. We instead formulate evaluation as a sequential allocation problem. Given a fixed trial budget and a set of scenarios whose failure behavior is unknown, which scenarios should be run, and run again? We propose a risk-aware contextual Thompson Sampling policy that combines a pre-execution scenario context vector and a fixed impact score with the failure outcomes observed during evaluation, and we test it by offline replay over 70 $τ$-bench airline scenarios and 824 recorded trials. Our main result is at the smallest budget: with only 50 trials ($6\%$ of the corpus), the policy recovers $86\%$ of the impact-weighted failures an oracle could find, compared to $25\%$ for uniform allocation. It discovers $3.5\times$ more impact-weighted failures (215.4 vs. 62.2) with the same number of trials, delivers $5\times$ the discovery per dollar, and cuts the budget wasted on scenarios that never fail from $34\%$ to $2.8\%$. The rest of our analysis demonstrates and qualifies this result: a budget sweep shows the advantage shrinks as the budget approaches the corpus size, and paired significance tests show that scenario context helps mainly at small budgets while posterior-based exploration helps at moderate ones. Risk-aware adaptive allocation therefore helps most exactly where evaluation budget is scarcest.