arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DynBranch:面向动态智能体LLM服务的推测性子图复用

DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving

Junyi Shen, Noppanat Wadlom, Zhengyuan Su, Yao Lu

arXiv 2609.31047首次发表:更新:

发表机构

National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DynBranch通过使未解决分支可寻址并复用子图结果,打破了智能体LLM服务中的分支解决障碍,在保持结果的同时大幅降低延迟。

AI 中文摘要

智能体LLM工作流在运行时决定其执行路径。下游计算可能是可预测的,或者之前已经运行过,但在模型或用户解决分支之前,它无法开始。我们将这种串行化称为分支解决障碍。仅靠缓存无法隐藏它:标识可复用结果的键直到那时才可知。在本文中,我们提出DynBranch,它使未解决的分支在解决之前变得可寻址。其稳定坐标允许候选子图在解决期间运行,并且完成的子图结果可以在后续请求中复用。一个两级控制器在预期收益超过负载代价时接纳此工作。DynBranch位于模型-API边界,不需要更改智能体框架或模型执行引擎。在四个智能体工作负载上,使用Qwen3-32B在4x H200 GPU上,DynBranch将平均延迟比每个工作负载最强的先前系统降低高达32%,比无复用基线降低46-66%,同时保持工作流结果。该收益在骨干家族和商品级Qwen3-8B/RTX 4090部署中持续存在。

英文摘要

Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Caching alone does not hide it: the key that identifies a reusable result is not known until then. In this paper, we propose DynBranch, which makes an unresolved branch addressable before it resolves. Its stable coordinate lets candidate subgraphs run during resolution and completed subgraph results be reused across later requests. A two-level controller admits this work when its expected benefit exceeds the load price. DynBranch sits at the model-API boundary and requires no changes to agent harnesses or model execution engines. Across four agentic workloads with Qwen3-32B on 4x H200 GPUs, DynBranch reduces mean latency by up to 32% over each workload's strongest prior system and by 46-66% against a no-reuse floor, while preserving workflow results. The benefit persists across backbone families and on a commodity Qwen3-8B/RTX 4090 deployment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑