arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AppWorld-UL:用于工具使用的多样化智能体-用户交互基准测试

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

Junzhi Chen, Harsh Trivedi, Jane Pan, Michael JQ Zhang, Tejas Srinivasan, Niranjan Balasubramanian, Ashish Sabharwal

arXiv 2607.20536首次发表:更新:

AI 中文总结

研究针对当前智能体-用户交互基准测试不足,引入 AppWorld-UL 基准测试,基于 AppWorld 框架及模拟应用构建多样任务,用大型语言模型模拟用户行为,评估显示其难度大,能推动回路中工具使用智能体研究。

AI 中文摘要

处理诸如订购杂货等日常数字任务的工具使用智能体不仅要操作应用程序,还需与用户交互。然而,当前评估智能体-用户交互的基准测试无法捕捉此类交互的多样性,且运行在小环境中,API 少且通常不改变状态。为此,我们引入了 AppWorld-UL,这是一个“用户在回路”的基准测试,包含 516 个具有挑战性的任务,要求多样化的智能体-用户交互。基于 AppWorld 框架及 9 个流行模拟应用构建,系统修改原始任务以引入模糊性和约束,促使各种类型的智能体-用户交互。用大型语言模型模拟用户行为,给出更可靠模拟。评估显示,最先进的大型语言模型 Claude Opus 4.7 在 AppWorld-UL 上成功率仅 48.6%,在更难的组合子集中仅 35.7%。在更严格的场景级指标上,组合任务性能降至 21.3%。分析表明正确的用户交互对成功至关重要,证明了该基准测试的难度及其推动回路中工具使用智能体研究的潜力。

英文摘要

Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for confirmation, and inform the user when the instruction is infeasible. However, current benchmarks for evaluating agent-user interactions do not capture the diversity of such interactions. Further, they operate in small environments with few, often non-state-changing, APIs. To address this gap, we introduce AppWorld-UL, a ``user-in-the-loop'' benchmark of 516 challenging tasks requiring diverse agent-user interactions. Building upon the AppWorld framework with 9 popular simulated apps like Amazon and Spotify, we systematically modify original tasks to introduce ambiguities and constraints that necessitate various types of agent-user interaction. User behavior is simulated by an LLM prompted to respond with carefully designed knowledge boundaries, offering more reliable simulation than the unconstrained or overly rigid alternatives used in prior work. Our evaluation reveals that a state-of-the-art LLM, Claude Opus 4.7, achieves only 48.6% success on AppWorld-UL, and only 35.7% on the harder, compositional subset. On the stricter, scenario-level metric, compositional task performance drops to only 21.3%. Our analysis reveals that correct user-interaction is crucial for success. This demonstrates the benchmark's difficulty and its potential to advance research on user-in-the-loop tool-use agents.

CommentsICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑