发表机构
Intuit Credit Karma(Intuit Credit Karma)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估协议保持上下文裁剪对智能体工作流的影响,发现协议感知裁剪和自适应预算护栏能显著提升任务成功率与协议遵循率,同时节省令牌,表明保留协议关键状态比最大化令牌移除更重要。
AI 中文摘要
智能体大语言模型(LLM)系统依赖长交互历史来保留指令、工具状态、中间决策和未解决的依赖关系,但无限制的上下文增长会增加计算成本并可能降低效率。本研究评估了协议保持上下文裁剪作为一种面向多步骤智能体工作流的可靠性约束方法。在保留上下文级别和工作流复杂度类别上,比较了五种裁剪策略——基于最近性、基于相关性、摘要、协议感知裁剪和自适应预算护栏——使用任务成功率、协议遵循率、有效工具调用、令牌节省、延迟降低、级联失败和关键上下文阈值等指标。传统策略实现了约60%的平均令牌节省,但任务成功率(66.6%-77.3%)和协议遵循率(85.5%-88.6%)较低。协议感知裁剪将任务成功率提升至92.2%,而自适应护栏实现了96.0%的任务成功率、96.3%的协议遵循率和1.0%的级联失败率,平均令牌节省56.0%。保留上下文预算为25%或更低时,失败几率相对于50%或更高的预算增加了10.92倍(p < 0.001)。在激进预算下,协议感知裁剪相比传统方法使成功完成几率提高5.24倍,而自适应护栏相比固定协议感知裁剪进一步将成功几率提高2.11倍(p < 0.001)。关键上下文阈值也随工作流复杂度增加而提高。这些发现表明,可靠的上下文缩减更多依赖于保留协议关键状态而非最大化令牌移除,且自适应护栏能在长时程智能体系统中提高效率、可扩展性和可靠性。
英文摘要
Agentic large language model (LLM) systems rely on long interaction histories to preserve instructions, tool states, intermediate decisions, and unresolved dependencies, but unrestricted context growth increases computational cost and can reduce efficiency. This study evaluates protocol-preserving context trimming as a reliability-constrained approach for multi-step agentic workflows. Five trimming strategies - recency-based, relevance-based, summarization, protocol-aware trimming, and adaptive budget guardrails - were compared across retained-context levels and workflow-complexity classes using task success, protocol adherence, valid tool calls, token savings, latency reduction, cascading failures, and critical context thresholds. Conventional strategies achieved about 60% mean token savings but lower task success (66.6-77.3%) and protocol adherence (85.5-88.6%). Protocol-aware trimming improved task success to 92.2%, while adaptive guardrails achieved 96.0% task success, 96.3% protocol adherence, and 1.0% cascading failure with 56.0% mean token savings. Retained-context budgets of 25% or less increased failure odds 10.92-fold relative to budgets of 50% or more (p < 0.001). Protocol-aware trimming produced 5.24-fold greater odds of successful completion than conventional methods under aggressive budgets, while adaptive guardrails further increased success odds 2.11-fold versus fixed protocol-aware trimming (p < 0.001). Critical context thresholds also increased with workflow complexity. These findings indicate that reliable context reduction depends more on preserving protocol-critical state than on maximizing token removal, and that adaptive guardrails can improve efficiency, scalability, and reliability in long-horizon agentic systems.
Comments13 Pages, 4 Figures