发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Long-Transduction诊断方法,测试长时程智能体在长时间生成中保持任务的能力,发现上下文长度、输入格式和任务复杂度分别导致62.8%、36.5%和39.9%的性能下降,揭示了关键可靠性风险。
AI 中文摘要
长时程智能体工作流要求模型在上下文不断增长、子任务复杂度变化以及新数据到达的同时,持续执行依赖于状态的动作。每种情形都代表智能体可能失败的一个独立维度。例如,一个核对长账本的智能体必须反复读取其状态、更新正确的记录,并在数千个输出中保持一致性。模型可能接受整个账本,但随着生成的进行,它可能丢失位置或停止一致地应用操作。我们引入了Long-Transduction,一种受控的诊断方法,用于测试模型在长时间生成过程中持续读取、修改和输出依赖于输入上下文的操作(如算术、排序、变量查找和表格转换)时保持任务的能力。Long-Transduction评估独立地变化局部任务复杂度、输入数据格式和上下文长度,以隔离每个维度上的失败。我们评估了七个开放权重模型,发现将上下文长度从4K扩展到128K时性能下降62.8%,改变输入格式时下降36.5%,增加局部任务复杂度时下降39.9%。这些失败共同代表了长时程智能体工作流中的关键风险。
英文摘要
Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an agent may fail. An agent reconciling a long ledger, for example, must repeatedly read its state, update the correct record, and preserve alignment across thousands of outputs. A model may accept the entire ledger yet lose its place or stop applying the operation consistently as generation proceeds. We introduce Long-Transduction, a controlled diagnostic that tests a model's ability to stay on task during long generation while continuously reading, mutating, and outputting input-context dependent operations such as arithmetic, sorting, variable lookups, and table transformations. Long-Transduction evaluation independently varies local task complexity, input data formatting, and context length isolate failures along each axis. We evaluate seven open-weight models, finding a 62.8\% decrease when scaling context length from 4-128K, a 36.5\% decrease when varying input format, and a 39.9\% decrease by increasing local task complexity. Together, these failures represent critical liabilities in long-horizon agentic workflows.