arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23321cs.DCcs.AIcs.LG

生产扩散模型推理服务中LoRA适配器的共现模式

Co-occurrence Patterns of LoRA Adapters in Production Diffusion Model Inference Services

Tao Zhang, Bin Liao, Tao Zhou, Yanping Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文基于阿里生产数据集GenTD26,用图论框架刻画LoRA适配器共现网络,发现其稀疏性、重尾分布及基础模型驱动等规律,并提出top-k预加载策略,为缓存与调度优化提供数据依据。

中文摘要 AI 辅助

低秩适配(LoRA)已成为在云端服务大规模个性化大语言模型和扩散模型的关键技术。然而,在生产推理工作负载下,适配器的共现模式、资源争用关系及演化规律尚未得到系统或定量的研究。基于阿里巴巴的生产扩散模型推理数据集GenTD26,本文采用图论框架构建适配器共现网络,并从静态结构和动态演化两方面进行特征刻画。我们的主要发现如下:(1)共现网络极其稀疏,适配器使用频率遵循显著的重尾分布。(2)引入第一个适配器会带来66.1%的执行延迟开销,此后边际成本递减。(3)共现关系由基础模型驱动:在90.6%的多适配器请求中,所有适配器共享相同的主导基础模型;66.2%的显著共现边连接同模型适配器对;在85.8%的多适配器请求中,所有适配器对均形成显著共现边。(4)适配器生态系统呈现核心-边缘双极结构,模型层面的周Jaccard相似度为0.696,12小时窗口内前10个最热门模型的更替率为54.5%。基于这些发现,我们提出了一种基于top-k共现统计的预加载策略;离线实验表明,在k=3时,该策略覆盖了测试集81.0%的共现对,跨频率阈值和时间窗口的敏感性分析验证了结论的稳健性。这些结果为LoRA推理服务中的缓存预加载、自适应调度和GPU内存管理提供了数据驱动的依据。

英文摘要

Low-rank adaptation (LoRA) has become a key technology for serving large-scale personalized large language models and diffusion models in the cloud. However, the co-occurrence patterns, resource contention relationships, and evolutionary regularities of adapters under production inference workloads have not been systematically or quantitatively studied. Based on GenTD26, Alibaba's production diffusion model inference dataset, this paper adopts a graph-theoretic framework to construct an adapter co-occurrence network and conducts a characterization from both static structure and dynamic evolution. Our main findings are as follows. (1) The co-occurrence network is extremely sparse, and adapter usage frequency follows a significant heavy-tailed distribution. (2) Introducing the first adapter incurs a 66.1% execution-latency overhead, with diminishing marginal costs afterwards. (3) Co-occurrence relationships are driven by base models: in 90.6% of multi-adapter requests, all adapters share the same dominant base model; 66.2% of significant co-occurrence edges connect same-model adapter pairs; and in 85.8% of multi-adapter requests, all adapter pairs form significant co-occurrence edges. (4) The adapter ecosystem exhibits a core-periphery bipolar structure, with a weekly Jaccard similarity of 0.696 at the model level and a churn rate of 54.5% for the top-10 hottest models within a 12-hour window. Based on these findings, we propose a preloading strategy built on top-k co-occurrence statistics; offline experiments show that it covers 81.0% of test-set co-occurrence pairs at k=3, and sensitivity analyses across frequency thresholds and time windows verify the robustness of the conclusions. These results provide a data-driven basis for cache preloading, adaptive scheduling, and GPU memory management in LoRA inference services.

发表机构

  • Guizhou University of Traditional Chinese Medicine(贵州中医药大学)
  • Guizhou University of Finance and Economics(贵州财经大学)
  • Beijing Technology and Business University(北京工商大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑