arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00558cs.SEcs.AI

AiFlow:面向流式大语言模型应用的带界背压令牌原生反应式编排框架

AiFlow: Token-Native Reactive Orchestration with Bounded Backpressure for Streaming LLM Applications

Qunhui Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有流式LLM工作流编排的背压等管理缺陷,提出AiFlow令牌原生反应式编排模型,经实验验证可降低应用TTFPT并控制队列深度。

中文摘要 AI 辅助

大语言模型(LLM)应用日益作为结合检索、工具调用、安全过滤器及多智能体协调的流式工作流运行。尽管现有框架暴露了提供商增量,但工作流节点常将生成视为粗粒度请求-响应步骤,将队列管理、工作者分配、排序及背压留给临时回调代码处理。本文提出AiFlow,一种令牌原生反应式编排模型,将提供商增量标准化为通过有向流式图传播的类型化Context<T>事件。每个节点由节点守护者管理,该守护者声明并执行本地队列边界、工作者并发、排序、溢出策略、取消传播及重试规则。我们形式化了有界内存属性,提出从紧凑DSL和JSON图形式的编译方法,并提供类型安全、状态并发及注入兼容性的静态验证。受控微基准、捕获的DeepSeek追踪重放(30次运行)、描述性在线运行、LangGraph基线、流式RAG工作负载及Ollama本地后端检查表明,AiFlow不改变提供商侧模型首字符生成时间(TTFT),但相比聚合方法将应用首令牌生成时间(TTFPT)降低70.9%-94.7%,且将运行时拥有的队列深度保持在声明边界内(相比无界策略最大队列深度降低93.7%-96.5%)。补充制品包含脚本、原始追踪、机器可读表格、校验和及无API冒烟测试;公共实现可通过FIT框架仓库获取。

英文摘要

Large language model (LLM) applications increasingly operate as streaming workflows combining retrieval, tool calls, safety filters, and multi-agent coordination. Although contemporary frameworks expose provider deltas, workflow nodes often treat generation as coarse request-response steps, leaving queue management, worker allocation, ordering, and backpressure to ad hoc callback code. This paper presents AiFlow, a token-native reactive orchestration model that normalizes provider deltas into typed Context<T> events propagated through a directed streaming graph. Each node is managed by a Node Guardian that declares and enforces local queue bounds, worker concurrency, ordering, overflow policy, cancellation propagation, and retry discipline. We formalize the bounded-memory property, present the compilation from a compact DSL and JSON graph form, and provide static validation for type safety, state concurrency, and injection compatibility. Controlled microbenchmarks, captured DeepSeek trace replay (30 runs), descriptive online runs, LangGraph baselines, a streaming RAG workload, and an Ollama local-backend check show that AiFlow does not alter provider-side Model TTFT but reduces Application TTFPT by 70.9-94.7\% versus aggregation and keeps runtime-owned queue depth within declared bounds (93.7-96.5\% MaxQ reduction versus unbounded policies). The supplementary artifact contains scripts, raw traces, machine-readable tables, checksums, and an API-free smoke test; the public implementation is available through the FIT Framework repository.

发表机构

  • School of Software, Shanghai Jiao Tong University(上海交通大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑