FlashVector:用于分层模型服务栈优化的智能体
FlashVector: Agent for Hierarchical Model Serving Stack Optimization
浏览论文内容
中文总结 AI 辅助
FlashVector是一个智能体系统,通过可扩展框架跨层优化模型服务栈,在Unity广告平台实现高达2倍吞吐量提升和1.98倍延迟加速。
中文摘要 AI 辅助
模型服务是生产推荐系统中最大的成本驱动因素之一。最大化其吞吐量需要穿越一个深度分层的层次结构:GPU内核、机器学习框架计算图、模型服务器以及按需特征处理——每一层都需要专门的专业知识。这种跨层专业知识本质上难以获取,并且无法随持续增长和演变的工作负载进行扩展,导致显著的成本效率收益未能实现。尽管最近的AI智能体在独立的GPU内核优化方面展示了人类专家级别的效率,但服务栈其余部分的自动化调优和优化在很大程度上仍未得到探索。我们提出了FlashVector,一个在模型服务栈的所有层优化性能的智能体系统。关键贡献是一个可扩展的框架,将单一内核优化智能体范式推广到异构技术栈,并整体性地交付性能改进。在部署到Unity的Vector广告平台后,FlashVector在模型服务器上实现了高达2倍的吞吐量提升和高达1.98倍的延迟加速,在特征存储上实现了高达1.6倍的吞吐量提升。这些优化不仅在GPU内核和计算图层面被发现,还跨越了模型服务栈的其他组件,如模型服务器(NVIDIA Triton的C++代码库)和按需特征转换服务(Python代码库),展示了该框架对更复杂系统架构的可扩展性。
英文摘要
Model serving is one of the largest cost drivers in production recommender systems. Maximizing its throughput requires navigating a deeply layered hierarchy: GPU kernels, the ML framework computation graph, the model server, and on-demand feature processing -- each demanding specialized domain expertise. Such cross-layer expertise is inherently difficult to acquire, and does not scale with a workload that continuously grows and evolves, leaving significant cost efficiency gains unrealized. While recent AI agents have demonstrated human expert level efficiency in standalone GPU kernel optimization, automated tuning and optimization for the rest of the serving stack remain largely unexplored. We present FlashVector, an agentic system that optimizes performance across all layers of the model serving stack. The key contribution is an extensible framework to generalize the single kernel optimization agent paradigm to heterogeneous technical stacks, and to deliver performance improvements holistically. After deployment in Unity's Vector advertising platform, FlashVector achieved up to 2x throughput increase and up to 1.98x latency speedup on model server, and up to 1.6x throughput increase on feature store. These optimizations were discovered not only at the GPU kernel and computation graph levels, but also across the other components of the model serving stack, such as the model server (NVIDIA Triton's C++ codebase) and the on-demand feature transformation service (Python codebase), demonstrating the extensibility of the framework to more complex system architectures.
发表机构
- Stanford University(斯坦福大学)
- Unity Vector AI Team(Unity Vector AI 团队)
机构由 AI 辅助整理,请以论文原文为准。