arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36488cs.LG

AdaptArena:评估Web代理的测试时个性化

AdaptArena: Evaluating Test-Time Personalization of Web Agents

  • Mila - Quebec AI Institute(米拉-魁北克人工智能研究所)
  • McGill University(麦吉尔大学)
  • ServiceNow Research(ServiceNow研究院)
  • Université Laval(拉瓦尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Dongchan Shin, Xing Han Lù, Jiaqi Deng, Jay Gala, Tomás Vergara Browne, Jaewon Moon, Fengyuan Liu, Alexandre Drouin, Siva Reddy, Alexandre Lacoste

AI总结:

AdaptArena通过隐含偏好推断评估Web代理的测试时个性化,发现性能差距大,强调推断与动作接地是关键挑战。

AI中文摘要:

大型语言模型(LLM)代理在复杂的网页导航任务中表现出强大的性能,但在用户意图未明确指定且偏好异质的现实环境中仍然脆弱。在实践中,用户很少提供明确的个人资料,要求代理从隐含信号中推断潜在偏好。尽管这对部署至关重要,但现有基准在很大程度上未探索这一问题设置。为解决这一差距,我们引入了AdaptArena,一个通过隐含偏好推断评估Web代理测试时个性化的基准。AdaptArena包含480个任务,涵盖单偏好和双偏好场景。每个评估任务必须通过检索和利用最相关的历史用户轨迹来解决,该轨迹隐含地编码了目标偏好。此外,我们引入了AdaptiveAgent,一个基于检索的框架,用于标准化评估隐含偏好推断。实验揭示了显著的性能差距:能够访问真实偏好的oracle代理实现了82.92%的成功率,而使用我们框架评估的LLM代理最多达到15.62%。此外,我们发现正确推断用户偏好对于任务成功是必要的但不充分的,因为即使代理与目标偏好对齐,下游网页交互中的执行失败仍然是一个重大瓶颈。这些发现强调了隐含偏好推断和稳健的动作接地是部署可靠、面向用户的Web代理的关键挑战。我们发布了我们的代码:此https URL

英文摘要:

Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent preferences from implicit signals. Despite its importance for deployment, this problem setting is largely underexplored in existing benchmarks. To address this gap, we introduce AdaptArena, a benchmark for evaluating test-time personalization of web agents via implicit preference inference. AdaptArena consists of 480 tasks, featuring both single-preference and double-preference scenarios. Each evaluation task must be solved by retrieving and leveraging the most relevant historical user trajectory that implicitly encodes the target preference. In addition, we introduce AdaptiveAgent, a retrieval-based framework for standardized evaluation of implicit preference inference. Experiments reveal a substantial performance gap: while oracle agents with access to ground-truth preferences achieve an 82.92% success rate, the evaluated LLM agents using our framework reach at most 15.62%. Furthermore, we find that correctly inferring user preferences is necessary but not sufficient for task success, as execution failures in downstream web interactions remain a significant bottleneck even when agents align with the target preference. These findings highlight implicit preference inference and robust action grounding as key challenges for deploying reliable, user-facing web agents. We release our code: https://github.com/McGill-NLP/web-agents-test-time-adaptations

↑