用于统一且优化执行的AI查询编译
AI Query Compilation for Unified and Optimized Execution
浏览论文内容
中文总结 AI 辅助
该研究提出统一编译执行架构范式,将SQL与LLM整合为张量计算图消除PCIe瓶颈,在SemBench数据集的TPU实验获最高5.3倍延迟、9.8倍吞吐量加速,还给出相关研究路线图。
中文摘要 AI 辅助
在这篇愿景论文中,我们提出了一种新的架构范式,用于通过统一编译执行策略加速AI查询执行。通过将混合AI查询作为整体进行编译——将标准SQL关系结构和LLM推理层整合为单一、统一的张量计算图——我们完全消除了执行边界间的PCIe数据移动瓶颈,并实现了全局编译器优化和高效自动分片。我们在SemBench Reviews和Movies数据集上的选定及扩展AI查询上验证了该统一执行范式的可行性,在TPU上实现了最高5.3倍的延迟加速和9.8倍的吞吐量加速,并概述了实现这一愿景需应对的开放技术挑战的研究路线图。
英文摘要
In this vision paper, we propose a novel architectural paradigm for accelerated AI query execution via a unified compiled execution strategy. By compiling the hybrid AI Query as a whole -- integrating both standard SQL relational constructs and LLM inference layers into a single, unified tensor compute graph -- we completely alleviate PCIe data movement bottlenecks across execution boundaries and enable global compiler optimizations and efficient automatic sharding. We demonstrate the viability of this unified execution paradigm on select and extended AI queries on SemBench Reviews and Movies datasets, achieving up to 5.3x latency speedup and 9.8x throughput speedup on TPUs, and outline a research roadmap of open technical challenges to realize this vision.