arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LongHorizon-Harness:推进面向真实世界任务的长时程智能体

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Ziyu Ma, Hailang Huang, Shun Zou, Yong Wang, Shidong Yang, Yiming Hu, Fei Wei, XiangXiang Chu

arXiv 2608.01964首次发表:更新:

发表机构

Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出LongHorizon-Harness框架,通过MEA循环等设计解决长时程智能体的状态追踪与错误传播问题,在多个基准测试中显著提升了Qwen、Claude等模型的表现。

AI 中文摘要

大型语言模型(LLM)智能体越来越多地承担需要在多个相互依赖的步骤中进行持续推理、工具使用和修正的长时程任务。然而,现有的智能体框架将任务执行、任务状态和完成度评估维持在不断增长的上下文内,导致状态难以追踪,且错误的自我评估会传播到后续决策中。我们将长时程执行重新表述为任务状态管理问题,并提出LongHorizon-Harness,它将任务状态显式维护在执行之外,仅通过从环境中独立验证的事实来更新状态。其Manage-Execute-Audit(MEA)循环使用一个管理器维护任务状态并确定下一个子任务,一个上下文全新的执行器执行该子任务,以及一个只读审计员在下一轮之前验证生成的环境状态。轻量型AgentAdapter支持模型和框架后端的互换,无需修改它们的原生智能体循环。LongHorizon-Harness将Qwen~3.7-Plus在WeaveBench上的表现从51.8%提升至80.7%,在Terminal-Bench~2.1上从69.7%提升至77.2%,在OSWorld~2.0上从2.8%提升至8.3%;还将Claude Opus~4.7在OSWorld2.0子集上的表现从20.0%提升至34.3%,证明其在模型、框架和交互领域中均能带来稳定提升。

英文摘要

Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing context, making the state difficult to track and allowing incorrect self-assessments to propagate into later decisions. We reformulate long-horizon execution as a task-state management problem and propose LongHorizon-Harness, which maintains the task state explicitly outside execution and updates it only with facts independently verified from the environment. Its Manage-Execute-Audit(MEA) loop uses a manager to maintain the task state and determine the next subtask, a fresh-context executor to perform it, and a read-only auditor to verify the resulting environment state before the next round. A lightweight AgentAdapter supports interchangeable model and harness backends without modifying their native agent loops. LongHorizon-Harness improves Qwen~3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench~2.1, and from 2.8% to 8.3% on OSWorld~2.0. It also raises Claude Opus~4.7 from 20.0% to 34.3% on an OSWorld2.0 subset, demonstrating consistent gains across models, harnesses, and interaction domains.

Comments29 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑