PM-Bench:评估大语言模型智能体中的前瞻记忆
PM-Bench: Evaluating Prospective Memory in LLM Agents
浏览论文内容
中文总结 AI 辅助
研究针对智能体人工智能中的前瞻记忆挑战,引入PM-Bench基准,受认知科学启发评估大语言模型智能体相关能力。在模拟一周内比较八个先进模型,发现各模型在此基准测试中均具挑战性,且无单一改善策略占优,还发布该基准用于诊断与干预。
中文摘要 AI 辅助
在智能体人工智能中,一个重大挑战是前瞻记忆,即在其他活动进行时,在特定未来线索或状态下执行意图的能力。我们引入了PM-Bench,这是一个基于文本的基准,用于衡量现代大语言模型智能体中的前瞻记忆能力。受认知科学中虚拟周范式的启发,PM-Bench评估大语言模型智能体维持用户意图、执行延迟意图以及监测潜在环境变化的能力。在模拟的七天一周时间里,智能体必须在决定是否有任何延迟任务到期时继续进行正在进行的活动。我们在八种不同的智能体配置下,在PM-Bench上比较了八个最先进的大语言模型。PM-Bench在所有设置下都具有挑战性:在我们的评估中,最佳方法,即GPT-5.4智能体,F1分数仅达到65.1%。此外,没有一种改善前瞻记忆的单一策略在所有模型中占主导地位。我们发布PM-Bench作为一个可控测试平台,用于诊断这些失败情况,并开发支持可靠前瞻行为的训练或推理时干预措施。
英文摘要
A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing. We introduce PM-Bench, a text-based benchmark for measuring prospective memory capabilities in modern LLM agents. Inspired by the Virtual Week paradigm from cognitive science, PM-Bench evaluates how well LLM agents maintain user intentions, execute delayed intentions, and monitor latent environment changes. Over the course of a simulated seven-day week, agents must continue an ongoing activity while deciding whether any deferred task is due. We compare eight state-of-the-art LLMs on PM-Bench under eight different agent configurations. PM-Bench proves challenging across all settings: the best method, a GPT-5.4 agent, reaches only 65.1\% F1 score under our evaluation. Furthermore, no single strategy for improving prospective memory dominates across models. We release PM-Bench as a controlled testbed for diagnosing these failures and developing training or inference-time interventions that support reliable prospective behavior.
发表机构
- University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。