面向交互式智能体的选择感知压力测试
Selection-Aware Stress Testing for Interactive Agents
浏览论文内容
中文总结 AI 辅助
该研究针对智能体评估的选择偏差问题,提出SASST方法,通过发现任务学习重加权并在独立确认任务评估,经实验验证其有效性,发现工作流程优势在确认阶段可能消失。
中文摘要 AI 辅助
智能体评估常使用同一基准选择工作流程,再搜索其优势减弱的任务类型,导致两项结论均来自同一数据。我们提出选择感知语义压力测试(\text{SASST}),该方法从发现任务的预执行特征中学习任务重加权,并在独立的确认任务上评估相同的配对比较。该协议检查支持性与稳定性,对所有计划主张使用联合边界,且可返回弃权(不执行)的主张。我们在给定的集群假设下证明了条件渐近有效性。一项含40个集群的审计发现高斯覆盖不足,以及保守的Bonferroni t边界。在一项含480个回合的τ基准研究中,3.75点的发现增益在确认阶段消失。第二项模型研究同样未确认工作流程优势或稳定压力规则。
英文摘要
Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. We introduce Selection-Aware Semantic Stress Testing (\SASST{}), which learns a task reweighting from pre-execution features on discovery tasks and evaluates the same paired comparison on separate confirmation tasks. The protocol checks support and stability, uses joint bounds for all planned claims, and can return no claim. We prove conditional asymptotic validity under stated cluster assumptions. A forty-cluster audit finds Gaussian undercoverage and conservative Bonferroni $t$ bounds. In one 480-episode $τ$-bench study, a $3.75$ point discovery gain vanished on confirmation. A second-model study likewise confirmed neither a workflow benefit nor a stable stress rule.
发表机构
- Purdue University(普渡大学)
- University of California, Irvine(加利福尼亚大学欧文分校)
机构由 AI 辅助整理,请以论文原文为准。