arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34045cs.DCcs.LG

Kafila:在可信的异构商用机器集合上服务大型语言模型

Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines

Murtaza Rangwala, Richard O. Sinnott, Rajkumar Buyya

首次发表
浏览论文内容

中文总结 AI 辅助

Kafila通过NAT组环和精确划分,在可信异构设备上高效服务大型语言模型,显著缩短流水线瓶颈,提升吞吐量。

中文摘要 AI 辅助

研究小组或朋友圈的成员共同拥有几台消费级计算机,但没有一台大到足以运行功能强大的大型语言模型。现有系统在开放的群体中汇集此类容量,任何人都可以加入,而只接纳可信机器的群体无法使用这些系统。限制成员资格消除了它们所依赖的东西:群体将模型的每个部分保存在多个对等节点上,并绕过速度慢的节点。有界会话必须使用它接纳的每台设备。其流水线以接收到无法快速服务的份额的设备的速度前进,因此划分必须在服务开始之前正确。我们提出了Kafila,其协议从NAT后面组装一个环,优先选择直接路径,在遍历失败时中继,而其规划器测量每台设备的内存带宽、容量和可达性,为固定的环顺序精确划分模型,并将持有嵌入和输出投影的头部与划分一起放置,而不是事先放置。在三个机群中具有不同能力的机器上,从共享局域网到跨越两大洲的五台设备,Kafila将最慢的流水线阶段缩短了最多5.2倍,相比流水线并行性的均匀划分(如GPipe),以及最多3倍,相比个人设备推理的内存比例划分(如exo),在那些划分低于一半的情况下,保持75%至87%的已投入硬件工作,并服务一个统一划分根本无法放置在机群上的模型。这对用户的价值取决于一个令牌中有多少是计算而不是网络。在成员共享网络的情况下,相同的划分返回1.56倍的均匀划分吞吐量和1.25倍的内存比例划分吞吐量,并且在四个并发用户下,该领先优势扩大到3.2倍而不是消失,每个用户几乎以单用户速率得到服务。

英文摘要

Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only trusted machines cannot use. Bounding membership removes what they depend on: a swarm holds each part of the model on several peers and routes around a slow one. A bounded session must use every device it admits. Its pipeline advances at the pace of whichever device received a share it cannot serve quickly, so the division has to be right before serving begins. We propose Kafila, whose protocol assembles a ring from behind NATs, preferring direct paths and relaying where traversal fails, while its planner measures each device's memory bandwidth, capacity and reachability, divides the model exactly for a fixed ring order, and places the head, which holds the embedding and output projection, together with that division rather than beforehand. On machines with different capabilities across three fleets, from a shared LAN to five devices spanning two continents, Kafila shortens the slowest pipeline stage by up to $5.2\times$ against the even split of pipeline parallelism, as in GPipe, and up to $3\times$ against the memory-proportional split of personal-device inference, as in exo, keeps 75 to 87 per cent of the committed hardware doing work where those divisions fall below half, and serves a model no uniform split can place on the fleet at all. What that is worth to a user depends on how much of a token is computation rather than network. Where the members share a network the same division returns $1.56\times$ the throughput of a uniform split and $1.25\times$ of a memory-proportional one, and under four concurrent users that lead compounds to $3.2\times$ rather than fading, each user served at almost the rate of one.

发表机构

  • The University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑