arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

最近最少使用策略下的多轮大语言模型对话:平均场渐近分析与命中率近似

Multi-Turn LLM Conversations under the Least-Recently-Used Policy: Mean-Field Asymptotics and Hit Ratio Approximation

Heyuan Yao, Chutong Gao, Yuan Lyu, Izzy Grosof, David Simchi-Levi

arXiv 2609.02027首次发表:更新:

发表机构

Northwestern University; Purdue University; HKUST(西北大学; 普渡大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多轮LLM服务系统,基于LRU策略建立模型,通过平均场渐近分析推导命中率闭式极限,提出命中率估算器并经Qwen3-8B模型实验验证,为系统分析和内存规划提供理论与实用支撑。

AI 中文摘要

现代大语言模型(LLM)服务系统的主要工作负载已从单次LLM调用转向多轮对话,其中新响应基于所有前序轮次的完整对话历史生成。命中率,即直接从高带宽内存(HBM)中现有缓存访问的KV缓存平均占比,是决定系统性能的关键指标。由于KV缓存前缀随轮次增长,且部分因有限内存容量必须被淘汰,系统动态极为复杂,估算命中率是一项极具挑战性的任务。我们将该系统建模为最近最少使用(LRU)策略下的多轮对话模型,通过平均场渐近框架证明:当对话到达率与内存容量成比例趋于无穷时,命中率收敛于闭式极限。基于该极限的特征,我们进一步提出一种实用的命中率估算器,并通过在昇腾神经网络处理器(Ascend NPUs)上实现的Qwen3-8B模型开展真实LLM服务实验验证其准确性。研究结果为多轮LLM服务系统分析提供了理论基础,也为内存容量规划提供了实用指导。

英文摘要

The major workloads in modern large language model (LLM) serving systems have shifted from single-shot LLM calls to multi-turn conversations, where new responses are generated based on the whole conversation history across all previous turns. The hit ratio, i.e., the average fraction of KV caches accessed directly from existing caches stored in high-bandwidth memory (HBM), is hence a crucial metric that governs system performance. Estimating the hit ratio is a highly nontrivial task due to the complex system dynamics, where the KV cache prefixes grow with turns and some must be evicted due to finite memory capacity. We formulate the system as a multi-turn conversation model under the least-recently-used (LRU) policy. Through a mean-field asymptotic framework, we prove that as the conversation arrival rate and the memory capacity grow proportionally to infinity, the hit ratio converges to a closed-form limit. Based on the characterization of the limit, we further propose a practical hit ratio estimator, and validate its accuracy by real LLM serving experiments on the Qwen3-8B model implemented on Ascend NPUs. Our results provide a theoretical foundation for the analysis of multi-turn LLM serving systems and a practical guideline for memory capacity provisioning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑