发表机构
Tencent; University of Maryland, College Park; University of Georgia; University of Minnesota, Twin Cities; Indiana University; Lehigh University; National University of Singapore; The Hong Kong Polytechnic University(腾讯; 马里兰大学帕克分校; 佐治亚大学; 明尼苏达大学双城分校; 印第安纳大学; 里海大学; 新加坡国立大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有终端基准测试局限,引入长视野终端基准测试(Long-Horizon-Terminal-Bench),含46个长视野任务。通过分解为分级子任务提供密集中间奖励,评估15个前沿模型,揭示改进空间,分析失败模式,发布该基准测试助力长视野终端智能体发展。
AI 中文摘要
人工智能智能体已能自主完成简短、明确的任务。然而,现有的终端基准测试大多聚焦于几分钟内就能完成的简单问题,且仅通过最终结果评估。这种设置忽略了中间进展和部分解决方案,产生稀疏奖励信号,无法全面了解智能体能力。我们引入了长视野终端基准测试,它包含46个长视野任务,涵盖九个类别。每个任务采用终端基准测试风格设置,有参考解决方案或模拟引擎,并进一步分解为细粒度的分级子任务。这使得能有密集的中间奖励和部分分数,不仅能评估智能体是否达成最终目标,还能了解其在开放式工作流程中的进展程度。这些任务通常需要数百个情节以及数分钟到数小时的执行时间,强调长视野规划、长上下文管理和迭代调试。我们评估了15个前沿模型,发现智能体平均每个任务消耗990万个令牌,每次运行大约有231个情节和85.3分钟的执行时间,这使得长视野终端基准测试比之前基于终端的基准测试要求更高。即使是测试中最强的模型,在部分奖励阈值为0.95时,通过率为15.2%,在完美奖励阈值为1.0时,通过率为10.9%,而各模型的平均通过率在两个阈值下分别为4.3%和1.7%。这些结果表明仍有改进空间。我们进一步分析了失败模式和错误模式,并发布长视野终端基准测试以支持长视野终端智能体的未来发展。
英文摘要
AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.
Comments17 pages