发表机构
Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该论文提出时序残差概念,指出长时序基准需将全任务成功与短阶段基线预测比较,以解释长任务失败的原因。
AI 中文摘要
长时序基准测试常显示,随着任务变长,智能体的失败情况会增多。这一观察对部署很有用,但本身并不能解释失败发生的原因。更多阶段会给普通错误的累积创造更多机会;更长的任务也可能包含更难的单个决策,或随着对话历史、工具输出和环境变化的累积而变得更难。我们用轨迹诱导退化来指代最后一种可能性:早期执行使后续工作更难。当这种有害的累积具体是模型可见的文本时,通常被称为上下文腐烂。在这篇立场文件中,我们认为,要声称存在“长时序失败”,基准测试必须将实际的全任务成功与从短的单个阶段构建的基线预测进行比较。我们将该预测与实际成功之间的对数比率称为时序残差。该比较必须使用相同的智能体配置,并预先规定如何选择阶段、检查点、信息和预算。残差表明,完整的执行过程与所选基线存在差异;仍需要针对性的实验来解释其原因。
英文摘要
Long-horizon benchmarks often show that agents fail more as tasks become longer. This observation is useful for deployment, but it does not by itself explain why failure occurs. More stages create more opportunities for ordinary errors to compound; longer tasks may also contain harder individual decisions or become harder as conversation history, tool outputs, and environment changes accumulate. We use trajectory-induced degradation to mean this last possibility: earlier execution makes later work harder. When the harmful accumulation is specifically the text visible to the model, it is often called context rot. In this position paper, we argue that to claim a "long-horizon failure", benchmarks must compare actual full-task success against a baseline prediction built from short, individual stages. We call the log-ratio between this prediction and actual success the horizon residual. The comparison must use the same agent configuration and specify in advance how stages, checkpoints, information, and budgets will be chosen. The residual shows that the full rollout differs from the chosen baseline; targeted experiments are still needed to explain why.