arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ProEvent:面向主动智能体的以事件为中心的基准测试

ProEvent: An Event-centric Benchmark for Proactive Agents

Guanzhen Li, Liangming Pan, Leye Wang

arXiv 2607.17701首次发表:更新:

发表机构

School of Computing, Peking University(北京大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对主动智能体研究忽视事件中心协助及评估难的问题,引入ProEvent基准测试,基于即时通讯聊天评估智能体维护用户时间表能力,实验发现当前智能体存在过度反应和事件取消处理难等问题,揭示了大语言模型的根本局限。

AI 中文摘要

主动智能体需在无明确指令时感知环境来预测用户需求并提供自主协助,识别和跟踪用户即将发生的事件是其基本能力。现有主动智能体研究多忽视以事件为中心的协助,且主动协助的开放性给可靠评估带来挑战。为此,我们引入ProEvent,首个以事件为中心的基准测试,基于即时通讯聊天评估智能体主动维护用户时间表的能力。它提供综合且现实的聊天内容,从响应时间、单步和多步响应正确性等方面评估。实验表明当前智能体常过度反应且难以处理事件取消,定性分析揭示了当前大语言模型作为主动智能体的根本局限。

英文摘要

Proactive agents are expected to anticipate user needs and provide autonomous assistance by perceiving environmental context without explicit instructions. A fundamental capability of such agents is to identify and track users' upcoming events, enabling continuous and event-specific assistance. For example, by recording the time and location of a planned hike, an agent can deliver weather reminders in advance or provide navigation support before departure. However, existing works on proactive agents largely overlook event-centric assistance, and the open-ended nature of proactive assistance poses challenges for reliable evaluation. To bridge these gaps, we introduce ProEvent, the first event-centric benchmark designed to assess an agent's ability to proactively maintain a user's timetable based on ongoing instant messaging chats. ProEvent provides synthesized yet realistic chats that consider the dynamic interaction among users, concurrent chat threads, and noise in the real world, and evaluates proactive agents on response timing, single-step response correctness, and multi-step response correctness. Experiments on eight LLMs and pipelines reveal that current agents frequently overact and struggle with event cancellation. Notably, even GPT-5.1 only reacts correctly in 26.7% of scenarios. Further qualitative analysis reveals fundamental limitations of current LLMs as proactive agents, particularly in detecting implicit events and reasoning from the user's first-person perspective.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑