发表机构
University of Southern California; Contextual AI; Capital One; Georgia Tech University(南加州大学; Contextual AI公司; 第一资本金融公司; 佐治亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有用户-智能体协作基准无法捕捉偏好动态的问题,推出AcCoRD基准,评估前沿LLM处理偏好动态的能力,发现模型难应对交互中出现/演变的偏好,仅提示无法实现所需不确定性识别。
AI 中文摘要
在用户-智能体协作中,用户偏好极少是静态且预先完全明确的:偏好会在交互过程中形成、显现、调整和放松。现有的用户-智能体协作评估基准几乎只聚焦于解决未明确的偏好,无法捕捉现实交互中更丰富的动态特性。我们推出AcCoRD,这是一个要求智能体处理在线购物和旅行规划两个领域中多样用户偏好动态的用户-智能体协作基准。我们在两种提示策略下评估了五个前沿大语言模型(LLM):标准ReAct和一种提示模型识别并解决用户偏好歧义的不确定性引导变体。我们的结果显示,前沿模型能够处理未明确的偏好,但难以满足交互过程中出现或演变的偏好,且需要更复杂的不确定性建模。此外,仅靠提示无法引出所需的不确定性识别能力。我们发布AcCoRD,作为开发能应对现实用户偏好全部复杂性的智能体的资源。
英文摘要
User preferences in user-agent collaboration are rarely static and fully-specified upfront: preferences are formed, revealed, adjusted, and relaxed during interaction. Existing benchmarks for evaluating user-agent collaboration focus almost exclusively on resolving underspecified preferences, thereby failing to capture the richer dynamics of real-world interaction. We introduce AcCoRD, a user-agent collaboration benchmark requiring agents to handle diverse user preference dynamics in two domains: online shopping and travel planning. We evaluate five frontier LLMs under two prompting strategies: vanilla ReAct, and an uncertainty-guided variant that prompts models to identify and resolve ambiguity about user preferences. Our results reveal that frontier models can handle underspecification but struggle to satisfy preferences that emerge or evolve mid-interaction and require more sophisticated uncertainty modeling. Further, prompting alone fails to elicit the required uncertainty recognition. We release AcCoRD as a resource for developing agents that can navigate the full complexity of real-world user preferences.