arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25061cs.CLcs.AIcs.DBcs.LGcs.PL

DataKernelBench:大语言模型能否在GPU上优化数据库查询?

DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

Gokul Karthik Kumar, Yotam Perlitz, Corey Lammie, Andrea Giovannini, Katja Hose

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出DataKernelBench基准,评估LLM在GPU上优化数据库查询内核的能力,实验显示最强模型配置在TPC-H数据集上实现2.11倍加速,工作负载上下文对性能影响大于硬件上下文。

中文摘要 AI 辅助

GPU正日益加速数据库系统,但特定查询的峰值性能仍常依赖手写内核。现有大语言模型(LLM)内核基准测试聚焦于机器学习算子,未测试不规则、异构、数据移动密集型的数据库风格算子。本文引入DataKernelBench,它将SQL转换为经过验证的PyTorch TorchPlan程序,并评估LLM通过执行引导修复优化CUDA或Triton中核心张量受限代码片段或完整查询的能力。在配备H100 GPU的TPC-H SF10数据集上,针对10个专有和开放权重模型,最强的完整查询CUDA配置在全通过率下实现了2.11倍的加速。研究发现,性能更高的实现通常使用内核融合和执行策略变更,更强的模型从完整查询专业化中受益最大,且工作负载上下文比硬件上下文更重要。为处理大于GPU内存的数据,本文扩展TorchPlan,在配备4个H100 GPU的TPC-H SF100上通过Dask-cuDF实现按需分区加载,达到了2.54倍的加速。

英文摘要

GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested. We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair. Across ten proprietary and open-weight models on TPC-H SF10 with an H100 GPU, the strongest full-query CUDA configuration achieves $2.11\times$ speedup over the TorchPlan baseline at full pass rate. We find that higher-performing implementations commonly use kernel fusion and execution-strategy changes, stronger models benefit most from full-query specialization, and workload context matters more than hardware context. To handle data larger than GPU memory, we extend TorchPlan with Dask-cuDF for on-demand partition loading on TPC-H SF100 with four H100 GPUs, achieving $2.54\times$ speedup. Project page: https://kerneldf.github.io/datakernelbench

发表机构

  • IBM Research(IBM研究院)
  • TU Wien(维也纳技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑