arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

卡利普索:关系型大语言模型服务

Kalypso: Relational LLM Serving

Hojae Son, Md Ashraful Islam, Huy Gia Cao, Hui Guan, Marco Serafini

arXiv 2607.23815首次发表:更新:

发表机构

UMass Amherst(马萨诸塞大学阿默斯特分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何提升大语言模型服务语义查询效率,提出关系型大语言模型服务,通过跨语义运算符流水线执行复用KV缓存状态,卡利普索系统利用自适应算法执行查询计划,解决在线调度问题,相比基线系统大幅提高查询完成时间。

AI 中文摘要

大语言模型越来越多地被用作语义运算符来处理非结构化数据。现有语义查询处理系统调用以请求为中心的大语言模型服务系统,未利用查询计划,存在性能提升空间。本文引入关系型大语言模型服务,使大语言模型服务能感知语义查询结构,保留查询语义和输出准确性。关键在于跨语义运算符的流水线执行,可复用KV缓存状态。提出卡利普索系统,通过自适应、内存感知调度算法执行语义查询计划。该系统解决了在线调度问题,平衡上游并行性、下游进度和GPU利用率。评估显示,卡利普索比基线系统提高了查询完成时间,加速比高达4.57倍,证明查询感知的大语言模型服务可显著提高语义查询执行效率。

英文摘要

Large language models are increasingly used as semantic operators for filtering, extracting, ranking, joining, and transforming unstructured data. Existing semantic query processing systems invoke request-centric LLM serving systems that are unaware of the query plan, leaving substantial performance opportunities unused. This paper introduces relational LLM serving, an abstraction that makes LLM serving aware of semantic query structure while preserving query semantics and output accuracy. The key opportunity is pipelined execution across semantic operators: when intermediate tuples flow directly from one operator to the next, their KV-cache state can be reused instead of recomputed. We present Kalypso, a relational LLM serving system that exposes an API for semantic query plans and executes them using an adaptive, memory-aware scheduling algorithm. Kalypso addresses a new online scheduling problem in which pipelined operator execution is coupled with GPU memory pressure management to reuse KV-cache state in the serving engine before eviction. Its scheduler continuously adjusts memory allocations to balance upstream parallelism, downstream progress, and GPU utilization. Our evaluation shows that Kalypso improves query completion time over baselines using request-centric LLM serving, with speedups up to 4.57x across diverse workloads, demonstrating that query-aware LLM serving can substantially improve the efficiency of semantic query execution.

Comments14 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑