arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Talaria:面向千亿参数语言模型的会话感知无服务器服务

Talaria: Session-Aware Serverless Serving of Hundred-Billion-Parameter LLMs

Utopia Meng, Unicornt Zhao, Derek Li, Goalen Gao, Frank Du

arXiv 2607.17181首次发表:更新:

AI 中文总结

研究面向千亿参数语言模型的无服务器服务问题,提出会话感知的Talaria系统,通过联合放置和准入决策保证会话连续性,其路由器排序放置,软预留考虑返回,会话预填充等措施提升性能,实验显示相比其他调度器有显著加速。

AI 中文摘要

无服务器多模型语言模型系统在共享GPU池上复用流行度不均衡的模型目录,但通常独立调度每个请求。使用工具的智能体打破了这种抽象:一个会话在短工具间隙中反复调用语言模型,携带一个长的可复用键值前缀,并由会话完成时间(SCT)来评判。仅负载路由会将一个延续与它的模型和键值状态分离,而基于轮次的模型复用甚至会将一个正确放置的延续延迟到目标模型的下一个时隙。对于千亿参数模型,这两种失败的代价都特别高昂:它们的权重限制了驻留,而长上下文键值的重建或移动成本很高。我们提出了Talaria,一个会话感知的无服务器多模型服务系统,它将会话连续性作为联合放置和准入决策。其路由器根据模型驻留、键值局部性和实例压力对放置进行排序,而软预留则在最后一个服务实例的准入预算中考虑可能的返回。会话预填充(SP)在活动模型时隙关闭之前准入符合预算的延续。一个实例本地底层保持HBM地址稳定,保留主机可恢复的键值,并在模型切换时暂存权重。在一台单TP=8的服务器上,我们在三个模型上重放了30个SWE-Bench模型会话(960次调用),每个模型的总参数都超过1000亿。与一个在禁用了SP、主机键值恢复和D2D暂存的情况下其他方面相同的轮次调度器相比,Talaria将p50 SCT从1000秒缩短到189秒,将p95从2296秒缩短到867秒,加速了5.3倍和2.6倍。

英文摘要

Serverless multi-model LLM systems multiplex popularity-skewed model catalogs over shared GPU pools, yet typically schedule each request independently. Tool-using agents break this abstraction: a session repeatedly calls an LLM across short tool gaps, carries a long reusable KV prefix, and is judged by session completion time (SCT). Load-only routing can separate a continuation from both its model and KV state, while round-based model multiplexing can delay even a correctly placed continuation until the target model's next slot. Both failures are especially costly for hundred-billion-parameter models: their weights constrain residency, while long-context KV is expensive to reconstruct or move. We present Talaria, a session-aware serverless multi-model serving system that makes session continuity a joint placement-and-admission decision. Its router ranks placements by model residency, KV locality, and instance pressure, while soft reservations account for likely returns in the last serving instance's admission budget. Session-prefill (SP) admits budget-eligible continuations before the active model slot closes. An instance-local substrate keeps HBM addresses stable, preserves host-restorable KV, and stages weights across model switches. On a single TP=8 server, we replay 30 SWE-Bench model-sessions (960 calls) over three models, each with more than 100B total parameters. Against an otherwise identical round scheduler with SP, host-KV restoration, and D2D staging disabled, Talaria cuts p50 SCT from 1000 s to 189 s and p95 from 2296 s to 867 s, speedups of 5.3x and 2.6x.

Comments19 pages, 15 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑