arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17188cs.CLcs.AI

多智能体AI工作流中的令牌优化与上下文窗口管理

Token Optimization and Context Window Management in Multi-Agent AI Workflows

Dvir Shamay

AI总结:

本文提出多智能体AI工作流的令牌优化与上下文窗口管理框架,含六种模式,可缩短延迟、减少令牌用量;经2420次试验发现相关性对比上下文可提升模型准确率,为模型研究与生产实践提供可重复方法。

AI中文摘要:

多智能体AI工作流的限制因素不仅包括模型质量,还包括令牌成本、延迟以及上下文窗口质量。本文提出了一种面向从业者的令牌优化与上下文窗口管理框架,该框架基于内部生产仪表盘构建,此仪表盘可从会议、电子邮件和聊天中提取结构化工作项,并通过大语言模型(LLM)在各工作流间路由摘要。文中描述了六种模式:上下文分层、一次获取/本地处理架构、模式压缩提示、令牌感知回退链、语义缓存以及智能体间通信压缩。在生产环境中,这些模式将实测的冷加载延迟从约3.5至10.5分钟的操作基线缩短至61至116秒(六次计时运行),并估计减少了60%至70%的令牌用量。本文还报告了一项受控上下文组合研究:针对11种模型配置开展的2420次验证性试验,使用了661项匿名工作场所项目并对其相关性进行评分。在保持提示固定为10个项目的情况下,用同领域低相关性项目替换部分高相关性项目,相较于仅使用高相关性项目,可提升模型对目标项目的相关性评分一致性;我们将此称为相关性对比上下文。在所有11项配对分析中,50:50信号/噪声条件相较于100%条件,相关性准确率提升了+0.077(朴素95%置信区间[+0.056, +0.098],Cohen's d=0.49,Holm校正后p<0.001,n=220)。这些样本并非独立;按9个模型系列计算,该效应为+0.084(95%置信区间[+0.064, +0.103]),报告为语料库内描述性比较,而非总体推断。Fusion-of-N的后续研究发现,学习到的合成效果并未优于项目ID的机械集合并集。本文的贡献是在模型研究与生产智能体实践之间构建了一个可测量的工程层:提供可重复的模式与评估方法,以实现更快速、更廉价、更可靠的工作流。

英文摘要:

Multi-agent AI workflows are limited not only by model quality but by token cost, latency, and context-window quality. This paper presents a practitioner framework for token optimization and context-window management, grounded in an internal production dashboard that extracts structured work items from meetings, email, and chat with LLMs and routes summaries across workstreams. Six patterns are described: context stratification, fetch-once/process-locally architecture, schema-contracted prompts, token-aware fallback chains, semantic caching, and inter-agent communication compression. In production they cut measured cold-load latency to 61-116 seconds (six timed runs) from an operational baseline of roughly 3.5-10.5 minutes, with an estimated 60-70% token reduction. It also reports a controlled context-composition study: 2,420 confirmatory trials across 11 model configurations, using 661 anonymized workplace items scored for relevance. Holding the prompt at a fixed ten items, replacing some high-relevance items with same-domain low-relevance items improves the model's relevance-score concordance on the target items, versus high-relevance items only; we call this relevance-contrast context. In the all-11 paired analysis, the 50:50 signal/noise condition improved relevance accuracy by +0.077 over the 100% condition (naive 95% CI [+0.056, +0.098], Cohen's d = 0.49, Holm-adjusted p < .001, n = 220). These cells are not independent; by the nine model families the effect is +0.084 (95% interval [+0.064, +0.103]), reported as a within-corpus descriptive comparison, not a population inference. A Fusion-of-N follow-up found that learned synthesis did not beat the mechanical set union of item IDs. The contribution is a measured engineering layer between model research and production agent practice: repeatable patterns and evaluation methods for faster, cheaper, more reliable workflows.

补充信息

↑