JarvisBench:人类与智能体间的始终在线智能
JarvisBench: Always-on Intelligence Between Humans and Agents
浏览论文内容
中文总结 AI 辅助
JarvisBench针对人类与智能体间的双向注意力协调问题,构建含45个任务实例的基准,可评估中间件的双向协调能力,为智能体能力提升提供稳定评估目标。
中文摘要 AI 辅助
长程智能体可持续执行任务,但人类注意力始终是间歇性且稀缺的,这形成了双向协调问题:用户在后台继续工作时可能需要立即访问智能体,而智能体可能会遇到需要用户判断的重要决策,此时用户已停止监控执行过程。我们提出一种始终在线的注意力协调层——Jarvis(得名于《钢铁侠》中的虚构AI助手),用于调解此接口并在一个或多个工作智能体间分配人类注意力。我们引入JarvisBench来评估双向协调:一是中间件能否准确且及时地回答用户关于正在进行的工作的问题,二是中间件能否识别智能体何时需要用户判断、在恰当的时刻请求该判断并将其反馈以改善任务结果。JarvisBench包含45个智能体任务实例:20个单智能体任务和25个工作流,组成10个多智能体项目,任务覆盖19个领域,从2000多个公开候选任务中筛选并调整而来。关键在于,用户注意力的需求在执行过程中自然产生,而非来自初始提示中的明显遗漏。JarvisBench被设计为可与任意智能体运行时集成,无需修改其底层执行循环,其参考实现还提供全双工语音接口,允许用户自然地访问Jarvis,同时及时的注意力协调支持智能体在后台工作。通过将智能体执行与注意力协调分离,JarvisBench在智能体能力不断提升的过程中提供了稳定的评估目标。
英文摘要
Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This creates a bidirectional coordination problem: users may need immediate access to an agent while work continues in the background, whereas agents may encounter consequential decisions that require user judgment after the user has stopped monitoring execution. We posit an always-on attention-coordination layer---\textit{Jarvis}\footnote{Named after the fictional AI assistant in \textit{Iron Man}.}---that mediates this interface and allocates human attention across one or more working agents. We introduce \textit{JarvisBench} to evaluate both directions of this coordination: whether an intermediary can accurately and promptly answer user-initiated questions about ongoing work, and whether it can recognize when an agent requires user judgment, solicit that judgment at the right moment, and route it back to improve task outcomes. JarvisBench contains 45 agentic task instances: 20 single-agent tasks and 25 workstreams organized into 10 multi-agent projects. The tasks span 19 domains and were selected and adapted from more than 2,000 public candidates. Crucially, the need for user attention arises naturally during execution rather than from an obvious omission in the initial prompt. JarvisBench is designed to integrate with arbitrary agent runtimes without modifying their underlying execution loops. Our reference implementation further provides a full-duplex speech interface, allowing users to reach Jarvis naturally while timely attention coordination supports agents working in the background. By separating agent execution from attention coordination, JarvisBench provides a stable evaluation target as agent capabilities continue to improve.
发表机构
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。