APeB:大型语言模型智能体个性化能力基准测试
APeB: Benchmarking Personalization Ability of Large Language Model Agents
浏览论文内容
中文总结 AI 辅助
研究大型语言模型智能体在处理原始、未明确查询时的个性化问题,通过引入个性化产品搜索测试平台构建基准测试,评估发现模型处理明确查询较好,但早期查询有困难,简单方法能提升性能。
中文摘要 AI 辅助
由大型语言模型驱动的智能体在用户发出原始、未明确的查询时,在个性化方面存在困难。在这种情况下,智能体必须推断潜在意图,从嘈杂的交互历史中提取偏好,并在相互竞争的选项中进行选择。现有基准很少测试这种能力……
英文摘要
LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract preferences from noisy interaction histories, and select among competing alternatives. Existing benchmarks rarely test this capability, as they often rely on user-refined queries or simplified histories. We introduce personalized product search (PPS), a testbed for agentic personalization under raw queries and diverse histories. We construct Agent Personalized Benchmark (APeB) from action logs, pairing underspecified intents with rich histories and user-viewed candidate items. Evaluating state-of-the-art LLMs with multi-step agent workflows, we find that models handle explicit queries well but struggle with early-stage queries requiring intent and preference discovery. Rubric analysis attributes this gap mainly to ineffective history use. A simple history-aware query-refinement pipeline, VQRA, yields consistent gains, highlighting the need for dedicated history-utilization modules in personalized agents.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。