发表机构
Institute of Artificial Intelligence, China Telecom (TeleAI); Shanghai Jiao Tong University(中国电信人工智能研究院(TeleAI); 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出\ extsc{Revise}方法,用于智能体工作流的细粒度在线恢复,平衡正确性与效率,在减少模型调用、提升服务性能方面效果显著。
AI 中文摘要
智能体修订在并发执行期间暴露出根本性的正确性与效率权衡问题:丢弃正在进行的工作可保证最新版本的正确性,但会浪费可能仍有效的进度;而复用先前工作可保证效率,但存在将过时状态传播到输出和工具效果中的风险。现有恢复策略通过粗粒度策略不平衡地解决这一权衡:要么通过允许潜在过时工作继续来偏向效率,要么通过重启工作流或从最早冲突处重新计算线性后缀来偏向正确性,从而丢弃未受影响的进度。我们提出了\ extsc{Revise},这是一种用于结构化智能体工作流中细粒度恢复的有效性引导运行时。当修订到达时,\ extsc{Revise}首先将其增量与记录的数据和控制依赖关系相交,并将所得影响传播到部分执行的DAG中以识别受影响的工作。随后,它停止无效工作,保留最早冲突之外已建立有效性的进度,且仅重新计算受影响的区域。不完整的来源会保守地扩大恢复范围,而复用的结果在提交前会被重新验证。对真实编码智能体轨迹的分析显示存在在线恢复机会:118个会话在排队的后续消息交付前保留了可观察的工作;在167个重叠的助手响应中,排队到完成的重叠时间在p95时达到56.55秒。在300个具有挑战性的修订/提交执行中,\ extsc{Revise}与无过时输出或效果的最新版本预言机表现相当。在使用Qwen3-14B的未修改LangGraph和LLMCompiler应用中,与完全重启相比,它减少了40.6--56.0%的模型调用,与后缀重新计算相比减少了31.3--43.6%的模型调用。在服务压力下,它进一步将修订到正确完成的令牌减少了13.26%,并将SLO良好吞吐量提高了3.07--5.43%。
英文摘要
Agent revisions expose a fundamental correctness--efficiency trade-off during concurrent execution. Discarding ongoing work preserves latest-version correctness but wastes progress that may remain valid, whereas reusing prior work preserves efficiency but risks propagating stale state into outputs and tool effects. Existing recovery strategies resolve this trade-off in an imbalanced way with coarse-grained policies: they either favor efficiency by allowing potentially stale work to continue, or favor correctness by restarting the workflow or recomputing a linear suffix from the earliest conflict, thereby discarding unaffected progress. We present \textsc{Revise}, a validity-guided runtime for fine-grained recovery in structured agent workflows. When a revision arrives, \textsc{Revise} first intersects its delta with recorded data and control dependencies and propagates the resulting impact through the partially executed DAG to identify affected work. It then stops invalid work, preserves validity-established progress beyond the earliest conflict, and recomputes only the affected region. Incomplete provenance conservatively expands recovery, while reused results are revalidated before commit. Analysis of real coding-agent traces show online recovery opportunities: 118 sessions retain observable work before a queued later message is delivered; across 167 overlapping assistant responses, enqueue-to-completion overlap reaches 56.55~s at p95. Across 300 challenging revision/commit executions, \textsc{Revise} matches a latest-version oracle with no stale outputs or effects. On unmodified LangGraph and LLMCompiler applications using Qwen3-14B, it reduces model calls by 40.6--56.0\% relative to full restart and by 31.3--43.6\% relative to suffix recomputation. Under serving pressure, it further reduces revision-to-correct-completion tokens by 13.26\% and improves SLO goodput by 3.07--5.43\%.