arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视界差距:长视界大语言模型智能体的规划、记忆、执行、训练与评估

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

Mingguang Chen, Licheng Wang, Bo Qu

arXiv 2608.06663首次发表:更新:

发表机构

DeepGrounding; AlphaAvatar(深地研究机构; 阿尔法 Avatar 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出长视界大语言模型智能体存在的视界差距,调研1547篇相关论文,厘清长视界等属性,分析领域应对措施并提出开放测量问题。

AI 中文摘要

前沿语言模型可通过单次前向传播解决多年前曾是研究成果的推理问题,但在耗时数小时的任务中表现不佳:会丢失之前的决策、宣称未完成的工作已完成,或偏离目标。我们将此称为视界差距,并通过系统种子采集(采用公开的26.8% bleed过滤器)及针对性补充,调研了2024至2026年间的1547篇arXiv论文。我们厘清了三个常被混淆的属性:长视界(任务属性:所需步骤)、长上下文(模型属性:token容量)和长期记忆(系统属性:跨步骤/会话的持久性)。我们将语料库按长视界任务的生命周期组织为六个类别——规划、记忆、执行、训练、评估,以及基础/安全,同时结合视界承载位置的维度(上下文内、任务内超出上下文、跨任务持久)。在所有类别中,我们发现相同模式:仅依赖结果的信号会随视界延长而变得无用,该领域的应对措施——无论是过程奖励模型、信用分配还是轨迹级诊断——都在生成更密集的步骤级信号。我们始终将批判性和诊断性文献视为核心线索,认为将评论与方法分离通常会导致单篇论文被拆分到不同章节。最后,我们提出了开放的测量问题:分解模型与工具的能力、管理同时用于训练和评估的过程级信号中的相关偏差,以及长视界可靠性是否存在通用预测理论。

英文摘要

Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done, or drifting from goals. We call this the horizon gap and survey 1,547 arXiv papers (2024-2026) collected via systematic seed harvest with a disclosed 26.8% bleed filter, extended by targeted supplementation. We disambiguate three routinely conflated properties: long-horizon (task property: required steps), long-context (model property: token capacity), and long-term memory (system property: persistence across steps/sessions). We organize the corpus into six categories tracking a long-horizon task's lifecycle -- planning, memory, execution, training, evaluation, and foundations/safety -- crossed with an axis capturing where horizons are carried (within-context, within-task-beyond-context, or cross-task-persistent). Across all categories, we find the same pattern: outcome-only signals grow uninformative as horizons lengthen, and the field's response -- whether process reward models, credit assignment, or trajectory-level diagnostics -- manufactures denser step-level signals. We treat critical and diagnostic literature as first-class threads throughout, arguing that segregating critique from method would routinely split single papers across chapters. We close by naming open measurement problems: decomposing model versus harness capability, managing correlated bias in process-level signals used for both training and evaluation, and whether long-horizon reliability admits general predictive theory.

Comments39 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑