arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

坚守任务:检验长时程智能体可靠性的基础

Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability

Jeffrey Willette, Krishna C. Puvvada, Boris Ginsburg

arXiv 2609.38712首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Long-Transduction诊断方法,测试长时程智能体在长时间生成中保持任务的能力,发现上下文长度、输入格式和任务复杂度分别导致62.8%、36.5%和39.9%的性能下降,揭示了关键可靠性风险。

AI 中文摘要

长时程智能体工作流要求模型在上下文不断增长、子任务复杂度变化以及新数据到达的同时,持续执行依赖于状态的动作。每种情形都代表智能体可能失败的一个独立维度。例如,一个核对长账本的智能体必须反复读取其状态、更新正确的记录,并在数千个输出中保持一致性。模型可能接受整个账本,但随着生成的进行,它可能丢失位置或停止一致地应用操作。我们引入了Long-Transduction,一种受控的诊断方法,用于测试模型在长时间生成过程中持续读取、修改和输出依赖于输入上下文的操作(如算术、排序、变量查找和表格转换)时保持任务的能力。Long-Transduction评估独立地变化局部任务复杂度、输入数据格式和上下文长度,以隔离每个维度上的失败。我们评估了七个开放权重模型,发现将上下文长度从4K扩展到128K时性能下降62.8%,改变输入格式时下降36.5%,增加局部任务复杂度时下降39.9%。这些失败共同代表了长时程智能体工作流中的关键风险。

英文摘要

Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an agent may fail. An agent reconciling a long ledger, for example, must repeatedly read its state, update the correct record, and preserve alignment across thousands of outputs. A model may accept the entire ledger yet lose its place or stop applying the operation consistently as generation proceeds. We introduce Long-Transduction, a controlled diagnostic that tests a model's ability to stay on task during long generation while continuously reading, mutating, and outputting input-context dependent operations such as arithmetic, sorting, variable lookups, and table transformations. Long-Transduction evaluation independently varies local task complexity, input data formatting, and context length isolate failures along each axis. We evaluate seven open-weight models, finding a 62.8\% decrease when scaling context length from 4-128K, a 36.5\% decrease when varying input format, and a 39.9\% decrease by increasing local task complexity. Together, these failures represent critical liabilities in long-horizon agentic workflows.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑