发表机构
Microsoft AI(微软人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过9个模型的大规模受控研究,发现LLM智能体在长时程生产决策任务中退化遵循几何规律,由步骤数而非上下文长度驱动,基准与生产条件间存在显著成功率差距,主张采用感知时程的评估与可靠性预算。
AI 中文摘要
尽管基准测试的成功率稳步提升,但大语言模型(LLM)智能体在生产环境中的长多步骤工作流部署仍不可靠。本文认为该差距主要是任务时程的产物:基准测试以中短时长任务为主,其成功率仍较高,而生产工作流所需的依赖步骤数量是基准的一个数量级。我们直接测量该效应,通过一项涵盖9个模型的大型受控研究,刻画智能体退化的形态并解析其成因,研究涉及6个参数规模从12亿到6710亿的开源模型、3个已部署的专有系统;4类任务家族,包括真正的智能体工具使用循环;5种时程;以及3种上下文 regime( regime 保留原词,指特定运行场景)。任务成功率遵循由单步可靠性参数决定的几何规律,该参数随模型规模提升,但即使是最强模型也远未达到1,保证了在足够长时程下的最终失效。该效应在智能体任务上最为显著,所有测试模型,包括广泛部署的系统,在16步内从接近完美的成功率降至接近零(分析轨迹数n=10664)。退化由步骤数而非上下文长度驱动:限定上下文窗口会加剧衰减而非缓解(logit斜率为-0.69,对比-0.44,p=3×10^-6),这与“中间丢失”解释相矛盾,并警示要警惕常见的生产捷径。将测得的可靠性投射到代表性基准时程上,量化出基准与生产条件间存在巨大差距:在GAIA长度的时程上为0.42,在百步骤生产时程上为0.24。对于负责大规模智能体编排与可靠性的团队,这些结果主张采用感知时程的评估与可靠性预算,而非依赖 aggregate pass-rate( aggregate 保留原词,指汇总的)指标。代码、提示词、随机种子及原始轨迹均已公开。
英文摘要
Production deployments of large language model (LLM) agents remain unreliable on long, multi-step workflows even as benchmark success rates climb steadily. We argue this gap is largely an artifact of task horizon: benchmarks are dominated by short-to-medium horizons where success remains high, while production workloads demand an order of magnitude more dependent steps. We measure the effect directly, characterizing the shape of agent degradation and disentangling its cause across a large controlled study spanning nine models, six open models from 1.2B to 671B parameters, and three deployed proprietary systems; four task families, including a genuinely agentic tool-use loop; five horizons; and three context regimes. Task success follows a geometric law governed by a single per-step reliability parameter, which rises with model scale but saturates well below 1 even for the strongest models, guaranteeing eventual collapse at sufficiently long horizons. The effect is sharpest on the agentic task, where every model tested, including widely deployed systems, falls from near-perfect success to near zero within sixteen steps of (n=10,664 analyzed trajectories. Degradation is driven by step count rather than context length: bounding the context window steepens decay rather than easing it (logit slope -0.69 vs. -0.44), p=3x10-6), contradicting a lost-in-the-middle explanation and warning against a common production shortcut. Projecting measured reliability onto representative benchmark horizons quantifies a substantial gap between benchmark and production conditions, from 0.42 at GAIA-length horizons to 0.24 at hundred-step production horizons. For teams responsible for agent orchestration and reliability at scale, these results argue for horizon-aware evaluation and reliability budgeting in place of aggregate pass-rate metrics. Code, prompts, seeds, and raw trajectories are released.