arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37267cs.AI

主动智能体的基础:原则、技术层次与 Proactivity-Gym

Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

  • KAIST(韩国科学技术院)
  • University of Minnesota(明尼苏达大学)

机构由 AI 辅助整理,请以论文原文为准。

Jio Oh, Seunghyun Do, Young-Jun Lee, Steven Euijong Whang, Dongyeop Kang

AI总结:

本文提出主动LLM智能体的3T原则(任务能力、时间分配、信任),构建五维设计空间及Proactivity-Gym测试平台,通过实验揭示性能差距与信任动态,强调三者的联合优化。

AI中文摘要:

主动式大语言模型(LLM)智能体可以在用户提出请求之前,将空闲计算转化为有用的支持。然而,即使工作正确,也可能误读用户情境、带来审查成本或削弱信任。本研究围绕三个联合原则(3T)提出了设计、实现和评估主动式LLM智能体的基础:任务能力(Task Capability),即预测相关需求并正确执行有用工作;时间分配(Temporal Allocation),即根据资源可用性和结果需求时间分配计算;以及信任(Trust),即维持用户对智能体的信心和适当的依赖。我们将这些目标与一个围绕五个维度组织的设计空间联系起来:任务范围、预测视野、激活触发、处理时机和干预深度,并指定了支持其选择所需的情境和系统建模,包括用户和环境表示、骨干LLM和智能体框架。最后,我们提出了PROACTIVITY-GYM,一个基于模拟的评估测试平台,包括多日场景、有状态环境和基于人格条件的模拟用户,可以评估主动协助在交互中的后果。在23种模型-框架配置上的评估揭示了3T方面的显著性能差距,并表明LLM评判者经常混淆任务能力和信任。一项有30名参与者的人类研究证明了3T联合优化的重要性:参与者在干预错位后即使结果正确,信任也会急剧下降,并且偏好睡眠时间协助(即使不完美)以保持持续的专注。总之,这些发现支持通过联合考虑有用工作、计算分配和不断演变的用户信任来设计和评估主动智能体。

英文摘要:

Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users' confidence and appropriate reliance on the agent. We connect these objectives to a design space organized around five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specify the situation and system modeling needed to support its choices, including user and environment representations, backbone LLMs, and agent harnesses. Lastly, we propose PROACTIVITY-GYM, a simulation-based evaluation testbed including multi-day scenarios, stateful environments, and persona-conditioned simulated users that can evaluate the consequences of proactive assistance across interactions. Evaluations across 23 model-harness configurations uncover substantial performance gaps across 3T and reveal that LLM judges often conflate task capability and trust. A human study with 30 participants demonstrates the importance of the joint 3T optimization: participants show sharp trust declines after intervention misalignment despite correct outcomes, and prefer sleep-time assistance, even when imperfect, to preserve ongoing focus. Together, these findings support designing and evaluating proactive agents through the joint consideration of useful work, compute allocation, and evolving user trust.

↑