现在询问,以后使用:评估长期 LLM 代理中的主动性差距
Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
- Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
- Noumena AI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对长期 LLM 代理在跨会话中未能主动获取用户偏好而导致的主动性差距,提出 Ask-to-Remember (ATR) 基准 ATRBench,通过隐藏用户偏好作为真实值来量化该差距,并诊断出获取环节是瓶颈。
AI中文摘要:
一个长期存在的 LLM 代理(例如 OpenClaw)的价值在于它能够根据用户跨会话的偏好和约束采取行动,而不仅仅是当前请求。然而,如今的代理会保留用户主动提供的信息,但很少询问那些未说出口的内容,这导致了长期 LLM 代理中的主动性差距:代理无法对从未获取到的偏好采取行动。随着用户将更多事务委托给代理,这种差距的影响也在增长。我们将这一差距的一个具体、可控的部分分离出来,称为 Ask-to-Remember (ATR):代理决定是否现在询问一个可重用的用户偏好,该偏好当前任务不需要,但后续与同一用户的会话会用到。ATR 甚至难以评估:正确的问题是不确定的,其回报会延迟到可能永远不会出现的任务。据我们所知,ATRBench 是第一个 ATR 基准,它通过将每个用户的偏好固定为隐藏的真实值,使得该差距可测量,因此成功需要询问,而不是回忆。在八个前沿 LLM 代理中,默认设置的表现至少比获得相关偏好的 oracle 低 62 分,而提示改进效果甚微。诊断表明获取是瓶颈。ATRBench 揭示了当前代理中的这一主动性差距,并提供了用于弥合该差距的诊断测试平台。
英文摘要:
A long-lived LLM agent, such as OpenClaw, earns its value by acting on a user's preferences and constraints across sessions, not just the current request. Yet today's agents keep what a user volunteers but rarely ask for what stays unspoken, leaving a proactivity gap in long-lived LLM agents: an agent cannot act on a preference it never obtained. As users delegate more of their affairs to agents, the impact of this gap grows. We isolate one concrete, controllable slice of this gap as Ask-to-Remember (ATR): the agent decides whether to ask now for a reusable user preference that the current task does not need but a later session with the same user will. ATR is hard even to evaluate: the right question is underdetermined and its payoff deferred to tasks that may never arise. ATRBench, to the best of our knowledge the first ATR benchmark, makes it measurable by fixing each user's preferences as hidden ground truth, so success demands asking, not recall. Across eight frontier LLM agents, defaults fall at least 62 points below an oracle handed the relevant preference, and prompting closes little of it. Diagnostics identify acquisition as the bottleneck. ATRBench surfaces this proactivity gap in current agents and offers a diagnostic testbed for closing it.