arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12551cs.DCcs.AI

RoofLang:实现AI驱动的LLM推理系统架构设计

RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems

Ziyue Yang, Yuting Jiang, Lei Qu, Peng Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

RoofLang是一种领域特定语言,通过通用工作负载表示、可验证变更空间和实现无关评估器,实现AI驱动的LLM推理系统架构设计,并发现新架构以提升吞吐量和交互性。

中文摘要 AI 辅助

AI开始对LLM推理优化做出实质性贡献。现有的AI优化主要基于性能剖析。受剖析约束的反馈将搜索限制在现有软件栈的能力和性能范围内,从而无法识别出从根本上更优的LLM推理系统架构。为了实现AI驱动的LLM推理系统架构设计循环,我们认为需要一种通用的工作负载表示、一个可验证的变更空间以及一个与实现无关的评估器。我们提出了RoofLang领域特定语言(DSL),它提供了这些特性。在我们的评估中,RoofLang揭示出DeepSeek V4系列模型相较于其他代表性模型,其峰值解码吞吐量可高出3.5-39.5倍。这一差距与其总参数量不成比例,主要源于紧凑的KV缓存设计,该设计支持更大的批处理并减少内存流量。一个持久化的优化智能体进一步发现了若干新架构,在NVIDIA B300上将DeepSeek V4 Pro的吞吐量和交互性提升了6.23-50.1%。

英文摘要

AI is beginning to make substantive contributions to LLM inference optimization. Existing AI optimizations are predominantly profiling-based. Profiling-bound feedback confines the search to the capabilities and performance of an existing software stack, preventing a fundamentally better architecture of LLM inference systems from being identified. To enable the AI-driven LLM inference system architecting loop, we argue that a general workload representation, a verifiable mutation space, and an implementation-independent evaluator are required. We present the RoofLang domain-specific language (DSL) that provides these features. In our evaluation, RoofLang reveals that DeepSeek V4-series models could achieve 3.5-39.5$\times$ higher peak decode throughput than other representative models. This gap is disproportionate to their total parameter counts and arises largely from compact KV-cache designs that support larger batches and reduce memory traffic. A persistent optimizer agent further discovered several new architectures that improved both throughput and interactivity of DeepSeek V4 Pro on NVIDIA B300 by 6.23-50.1%.

发表机构

  • Shanghai Xingyunzhili Artificial Intelligence Institute(上海星云智理人工智能研究院)
  • Microsoft Research(微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑