发表机构
MATS; ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center(MATS; ELLIS 图宾根研究所; 马克斯·普朗克智能系统研究所; 图宾根人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgentTime基准测试智能体按请求时长工作、预测运行时间及事后估计已用时间的能力,发现完成任务不等于能控制自身时间,需同时评估两者。
AI 中文摘要
智能体的一项关键控制能力是管理自身运行时间。这种能力需要时间感知,以预测和估计墙钟时间并控制自身行为。先前工作聚焦于时间感知,但在原生智能体框架中的持续时间遵循与控制仍未探索。我们提出AgentTime,一个用于测试智能体能否按请求时长工作、预测自身运行时间以及事后估计已用时间的基准。它包含来自18个来源的222个任务,涵盖编码、计算机使用、智能体工作和自动化研究。持续时间遵循实验附加一条指定工作时长的指令,请求范围从约一分钟到多天。这些指令的准确性差异显著:Claude Code中的Fable 5.1偏离请求运行时间的典型因子为2.9倍,而Codex中GPT-6 Astra仅为1.2倍。然而,匹配请求运行时间本身并不确保持续工作于任务。在158个具有可分类转录的Astra运行中,14个在看似完成后明确休眠。在预测实验中,预测倾向于高估自然运行时间。在回顾实验中,移除时间信息使Sol和Astra的偏差增加一倍以上,Fable则接近一倍。智能体完成任务的能力并不保证其能控制自身时间或按整个请求时长工作。为使智能体在长时间跨度上可靠、安全且自主运行,我们需要对两者进行评估。
英文摘要
An essential control of AI agents is their ability to manage runtime. This ability requires a sense of time-awareness, to predict and estimate wall-clock time and to control their own actions. Prior work has focused on time-awareness, but duration-following and control in native agent harnesses remain unexplored. We present AgentTime, a benchmark for testing whether agents can work for a requested duration, predict their runtime, and estimate elapsed time afterward. It comprises 222 tasks from 18 sources spanning coding, computer use, agentic work, and automated research. Duration-following experiments append a single instruction specifying how long to work, with requests ranging from about a minute to multiple days. Accuracy on these instructions varies substantially: Fable 5.1 in Claude Code deviates from requested runtimes by a typical factor of 2.9$\times$, compared with only 1.2$\times$ for GPT-6 Astra in Codex. However, matching the requested runtime does not, by itself, establish continued work on the task. Among 158 reviewed Astra runs with classifiable transcripts, 14 explicitly slept after appearing to finish. In forecasting experiments, predictions tend to overestimate natural runtimes. In retrospective experiments, removing temporal information more than doubles deviation for Sol and Astra and nearly doubles it for Fable. An agent's ability to complete a task does not guarantee that it can control its own time or work for the whole requested duration. For agents to run reliably, safely, and autonomously over long horizons, we require the evaluation of both.
CommentsWebsite: https://agenttimebench.com Code: https://github.com/michaelofengenden/agenttimebench