arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PDEU-Bench:基准测试工具调用LLM智能体的个性化规划生命周期

PDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM Agents

Huayi Lai, Shichao Song, Qingchen Yu, Simin Niu, Mengwei Wang, Hanyu Wang, Xun Liang

arXiv 2609.34930首次发表:更新:

发表机构

Renmin University of China; Beihang University(中国人民大学; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对工具调用LLM智能体在长期交互中个性化规划能力评估缺失的问题,提出PDEU-Bench基准,覆盖214个任务和94个工具,发现模型在计划定义与更新上存在明显短板,现有方法难以全周期保持用户偏好。

AI 中文摘要

大型语言模型(LLM)智能体正从执行孤立指令的工具调用系统,演变为通过持续、多步骤交互来追求用户目标的任务导向型智能体。然而,现有的个性化工具使用基准大多评估孤立调用或反应式执行,尚不清楚智能体能否在长期交互中制定、执行并修订明确计划,同时保持用户偏好。为弥补这一空白,我们提出了PDEU-Bench(个性化计划定义、计划执行与计划更新基准),这是一个用于评估个性化工具使用智能体完整规划生命周期的基准。PDEU-Bench包含214个长时程交互任务,覆盖12个日常领域和94个工具,并针对偏好遵循和计划质量进行分阶段评估。对15个代表性开源和闭源LLM的广泛评估揭示了局部工具执行与动态规划之间的显著差距:LLM通常能在单个调用中实例化偏好,却难以构建连贯的计划定义和计划更新。我们进一步评估了主流个性化与记忆增强方法。尽管这些方法改善了特定阶段,但没有任何一种被评估的方法能在整个生命周期中可靠地传播用户偏好,且其收益往往无法转移到后续执行中。细粒度的错误分析进一步揭示,偏好遗漏和冲突在整个规划生命周期中持续存在,凸显了未来研究需以偏好感知的信息检索和记忆能力对LLM进行参数化的必要性。我们在附录中提供了相关代码和数据,以支持未来研究。

英文摘要

Large language model (LLM) agents are evolving from tool-calling systems that execute isolated instructions into task-oriented agents that pursue user goals through sustained, multi-step interactions. However, existing benchmarks for personalized tool use largely assess isolated calls or reactive execution, leaving unclear whether agents can formulate, execute, and revise an explicit plan while preserving user preferences throughout long-term interaction. To address this gap, we introduce \textbf{PDEU-Bench} (\textbf{P}ersonalized plan \textbf{D}efinition, plan \textbf{E}xecution, and plan \textbf{U}pdate \textbf{Bench}mark), a benchmark for evaluating the complete planning lifecycle of personalized tool-using agents. PDEU-Bench comprises 214 long-horizon interaction tasks spanning 12 everyday domains and 94 tools, with stage-specific assessments of preference adherence and plan quality. Extensive evaluations of 15 representative open-source and closed-source LLMs reveal a pronounced gap between local tool execution and dynamic planning: LLMs can often instantiate preferences in individual calls, yet struggle to construct coherent plan definition and plan update. We further evaluate mainstream personalization and memory-augmentation methods. Although these methods improve particular stages, none of the evaluated methods reliably propagates user preferences throughout the complete lifecycle, and their gains frequently fail to transfer to subsequent execution. Fine-grained error analysis further reveals that preference omissions and conflicts persist throughout the planning lifecycle, highlighting the need for future research to parameterize LLMs with preference-aware information retrieval and memory capabilities. We provide the relevant code and data in the appendix to support future research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑