arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

[AAFLOW+] 用于多智能体工作流的具有零拷贝分布式键值缓存编排的有状态算子抽象

[AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows

Arup Kumar Sarker, Alexander James Halpern, Mills Staylor, Aymen Alsaadi, Gregor von Laszewski, Yue Cheng, Shantenu Jha, Geoffrey Fox

arXiv 2607.10987首次发表:更新:

AI 中文总结

研究多智能体语言模型系统文本中心问题,提出AAFLOW+扩展,构建通信感知图并提供多种算子,实现零拷贝执行,通过分析成本模型验证,大幅降低TTFT、计算成本、键值内存并提高吞吐量,提升多智能体系统效率

AI 中文摘要

多智能体语言模型系统越来越多地集成检索、规划和推理,但仍以文本为中心,需要智能体通过昂贵的预填充来反复重新计算共享上下文。虽然已知单请求推理可通过键值缓存管理加速,但通常限于本地服务范围。我们引入了AAFLOW+,它是智能体工作流算子的有状态扩展,使键值缓存成为一等分布式系统对象。AAFLOW+将流程构建为通信感知图,同时优化数据、提示和可重用模型状态。它还提供了键值物化、传输、分叉、组合和逐出的算子。其运行时支持零拷贝、传输感知执行,允许智能体重用长上下文而无需重新计算。基于由经验硬件微基准参数化的分析成本模型,AAFLOW+将首次令牌到首次令牌时间(TTFT)降低了高达50.2倍,在16智能体规模下将多智能体计算成本降低了高达7.63倍,将键值内存减少了1.72 - 6.10倍,并将吞吐量提高了超过7.74倍。结果表明,在中等到高带宽的网络上,键值传输优于重新计算,通过取代文本传递确保键值状态共享大大提高了多智能体语言模型系统的效率。

英文摘要

Multi-agent LLM systems increasingly integrate retrieval, planning, and reasoning, but remain fundamentally text-centric, requiring agents to repeatedly recompute shared context through expensive prefill. Although single-request inference is known to be accelerated by KV-cache management, it is usually restricted to local serving scopes. We introduce AAFLOW+, a stateful extension of agentic workflow operators that makes KV cache a first-class distributed systems object. AAFLOW+ builds processes into communication-aware graphs that concurrently optimize data, prompts, and reusable model state. It also provides operators for KV materialization, transfer, fork, composition, and eviction. Its runtime enables zero-copy, transfer-aware execution, allowing agents to reuse long context without recomputation. AAFLOW+ reduces TTFT by up to 50.2x, achieves up to 7.63x reduced multi-agent compute cost at 16-agent scale, reduces KV memory by 1.72-6.10x, and increases throughput by more than 7.74x, based on an analytical cost model parameterized by empirical hardware microbenchmarks. The results demonstrate that KV transmission outperforms recomputation on networks with moderate to high bandwidth, making sure KV-state sharing greatly increases efficiency in multi-agent LLM systems by replacing text passing.

Comments21 pages, 10 Figures, 12 Tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑