arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08268cs.DCcs.AIcs.LGcs.SYeess.SY

DySCo:面向边缘-云协同LLM推理的动态分片与深度同步批处理

DySCo: Dynamic Sharding for Collaborative Edge-Cloud LLM Inference with Depth-Synchronized Batching

Jingpo Xu, Paul Joe Maliakel, Ivona Brandic, Shashikant Ilager

首次发表
浏览论文内容

中文总结 AI 辅助

DySCo通过动态分片和深度同步批处理,解决边缘-云LLM推理中的空闲间隙和异构批处理低效问题,显著提升吞吐量并降低延迟。

中文摘要 AI 辅助

智能应用日益广泛地部署在移动和物联网(IoT)边缘设备上。因此,大型语言模型(LLMs)越来越多地被用于支持这些应用。然而,由于其高资源需求,LLMs大多部署在云端。分层边缘-云推理使得资源受限的边缘设备能够为它们无法完整托管的LLMs贡献计算能力。然而,异构的分割点引入了两个耦合的低效问题。首先,边缘执行和通信在云端调用之间产生了空闲间隙。其次,到达不同模型深度的请求无法进行常规批处理。我们提出了DySCo,一个协作运行时,它保持KV缓存本地化,并引入了dyForward,一个模型感知的层范围执行器,能够从驻留的模型分片上运行可配置的连续层范围,而无需重新加载权重。对于多边缘服务场景,我们引入了深度同步批处理(DSB),它将异构请求推进到最深的切分点,并对其共同后缀进行批处理。跨异构设备、两个模型家族以及本地和广域网链路的实验表明,即使排除等待时间,空闲间隙也会增加后续GPU前向调用的延迟,在我们的测量中,每个解码步骤额外增加高达25毫秒的云端后缀延迟。在平均并发数为8时,DSB相比FIFO提高了275%的吞吐量,相比精确匹配批处理提高了48%,相比轮询交错提高了79%,同时降低了平均每会话延迟。这些结果表明,具有不同边缘-云分割的请求可以重用驻留的云端权重,并共享批处理的后缀计算。本工作的工件仓库公开可用,地址为:此https URL

英文摘要

Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasingly used to support these applications. Yet, due to their high resource demands, LLMs are mostly deployed in the cloud. Layer-wise edge-cloud inference lets resource-constrained edge devices contribute computation to LLMs they cannot host in full. However, heterogeneous split points introduce two coupled inefficiencies. First, edge execution and communication create idle gaps between cloud invocations. Second, requests arriving at different model depths cannot be conventionally batched. We present DySCo, a collaborative runtime that keeps KV caches local and introduces dyForward, a model-aware layer-range executor that runs configurable contiguous layer ranges from resident model shards without reloading weights. For multi-edge serving settings, we introduce depth-synchronized batching (DSB), which advances heterogeneous requests to the deepest cut and batches their common suffix. Experiments across heterogeneous devices, two model families, and local and wide-area links show that idle gaps increase the latency of subsequent GPU forward calls even when waiting time is excluded, adding up to 25 ms of additional cloud-side suffix latency per decoding step in our measurements. At an average concurrency of eight, DSB improves throughput by 275% over FIFO, 48% over exact-match batching, and 79% over round-robin interleaving while reducing mean per-session latency. Together, these results show that requests with different edge-cloud splits can reuse resident cloud weights and share batched suffix computation. The artifact repository for this work is publicly available at: https://github.com/Large-scale-Sustainable-Computing-LSC/dysco-artifact

发表机构

  • University of Amsterdam(阿姆斯特丹大学)
  • TU Wien(维也纳工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑