AI 中文总结
该研究针对LLM基础设施中张量管理策略复用受阻的问题,提出TensorCast分布式张量管理层,将张量状态管理与计算逻辑解耦,在保持性能的同时提升多轮智能体工作负载的TTFT达93.2%。
AI 中文摘要
现代大语言模型(LLM)基础设施日益将张量不仅作为计算数据进行管理,还作为跨分布式组件共享的持久状态进行管理。现有系统通过将特定任务机制与执行引擎、网络或存储后端深度集成,优化单个张量管理任务,例如模型权重加载、KV缓存管理和检查点同步。然而,这种专业化形成了孤立的竖井,阻碍了张量管理策略在不断发展的LLM工作负载中的复用和组合。在本文中,我们将张量生命周期管理确定为LLM基础设施中缺失的抽象层,并提出了张量即服务(TaaS),该服务将张量状态管理与计算逻辑解耦。我们设计并构建了TensorCast,这是一个分布式张量管理层,它提供一流的张量抽象、可编程的原语生命周期,以及将张量管理策略与执行机制分离的运行时。这使开发人员能够使用TensorCast API编写张量管理程序,同时透明地利用分布式执行和数据移动。我们将TensorCast与vLLM和SGLang集成,并在各种张量生命周期工作负载中对其进行评估,包括模型权重物化、权重同步、KV缓存管理和可编程请求路由。我们的结果表明,TensorCast在实现与专用张量管理系统相当的性能的同时,还支持新的跨组件优化策略。使用TensorCast实现的可编程策略在高度并发的多轮智能体工作负载下,将中位数首包生成时间(TTFT)提高了多达93.2%。
英文摘要
Modern LLM infrastructure increasingly manages tensors not only as computation data, but also as persistent states shared across distributed components. Existing systems optimize individual tensor management tasks, such as model weight loading, KV cache management, and checkpoint synchronization, by deeply integrating task-specific mechanisms with execution engines, networks, or storage backends. However, this specialization creates isolated silos that hinder the reuse and composition of tensor management strategies across evolving LLM workloads. In this paper, we identify tensor lifecycle management as a missing abstraction layer in LLM infrastructure and propose Tensor-as-a-Service (TaaS), which decouples tensor state management from computation logic. We design and build TensorCast, a distributed tensor management layer that provides first-class tensor abstractions, programmable lifecycle primitives, and a runtime that separates tensor management policies from execution mechanisms. This enables developers to write tensor management programs using TensorCast APIs while transparently leveraging distributed execution and data movement. We integrate TensorCast with vLLM and SGLang and evaluate it across diverse tensor lifecycle workloads, including model weight materialization, weight synchronization, KV cache management, and programmable request routing. Our results show that TensorCast achieves competitive performance with specialized tensor management systems while enabling new cross-component optimization policies. A programmable policy implemented with TensorCast improves median TTFT by up to 93.2% under highly concurrent multi-turn agent workloads.