发表机构
Carnegie Mellon University; Microsoft Research(卡内基梅隆大学; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CAVEAT基准,揭示计算机使用智能体在激励错位环境中易受引导机制影响,并开发CAVEAT-Harness干预方法,显著提升用户最优购买率。
AI 中文摘要
计算机使用智能体(CUA)越来越多地代表用户在线执行任务。当它们运行的环境存在与用户目标不一致的激励时会发生什么?例如,在在线市场中,平台可能偏向某些产品,从而可能使智能体偏离用户的目标。现有的CUA基准测试覆盖了合作性设置或显式攻击,但并未测试当环境本身在结果中有利害关系时,智能体是否保持用户目标。我们引入了CAVEAT,一个受控基准测试,涵盖九个市场环境以及八种常见引导机制的分类。在五个模型家族中,智能体在匹配对照情节中购买用户最优产品的比例为78.6%,但在启用引导机制时仅为17.3%。更大的模型和增强的推理能力提高了鲁棒性,但显著的失败仍然存在。我们的轨迹分析和针对性消融识别出引导进入决策过程的三个关键点:(1)智能体扭曲用户的优先级,(2)过早缩小其考虑的备选方案集,(3)在解决决策相关证据之前就做出承诺。基于这一诊断,我们开发了CAVEAT-Harness,它直接针对这些失败模式,将用户最优购买率提高了55.0%。针对性的后训练进一步改进了一个较小的开放模型。这些结果确立了激励鲁棒性作为委托智能体的一个独特挑战,诊断了其失败方式,并表明针对性干预可以显著改善它。
英文摘要
Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives of their own? Online marketplaces, for example, may favor some products over others, steering agents away from the user's objective. Existing CUA benchmarks cover cooperative settings or explicit attacks, but do not test whether agents preserve user objectives when the environment itself has a stake in the outcome. We introduce CAVEAT, a controlled benchmark spanning nine marketplace environments and a taxonomy of eight common steering mechanisms. Across five model families, agents purchase the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering mechanisms are enabled. Larger models and more reasoning improve robustness, but substantial failures persist. Trajectory analysis and targeted ablations identify three weaknesses in how agents decide: they (1) prematurely narrow the set of alternatives they consider, (2) impose priorities the user never stated, and (3) commit before resolving decision-relevant evidence. Guided by this diagnosis, we develop CAVEAT-Harness, which targets these failures and raises the optimal purchase rate by up to 80.0 percentage points, and show that targeted post-training further improves a smaller open model. These results establish incentive robustness as a distinct challenge for delegated agents, diagnose failure modes, and show how targeted interventions can substantially improve robustness.