展示GenDB:通过大语言模型代理实现实例优化和定制化查询处理代码生成
Demonstrating GenDB: Instance-Optimized and Customized Query Processing Code Generation via LLM Agents
浏览论文内容
中文总结 AI 辅助
研究针对传统查询处理引擎扩展难题,提出GenDB生成式查询引擎,借助大语言模型代理生成代码。核心方法是通过LLM agents为特定场景生成优化代码,主要贡献是展示其工作流程、性能优势并提供用户探索平台。
中文摘要 AI 辅助
传统查询处理引擎因内部复杂性难以扩展,构建新系统需大量工程投入。为解决此问题,本文展示了GenDB这一生成式查询引擎,它将查询处理从手动构建系统转变为由大语言模型驱动的查询处理代码生成。早期原型利用大语言模型代理生成针对特定数据、工作负载和硬件资源的实例优化查询执行代码,适用于重复性模板化查询的离线代码生成。对于即席查询,GenDB可与传统数据库管理系统以混合架构协作。本文演示让用户能可视化交互探索GenDB的工作流程,通过视觉检查和分析了解其性能优势,还能上传自己的数据和查询进行探索。
英文摘要
Traditional query processing engines require continuous development and extensions to support new techniques and user requirements, and in some cases, entirely new systems must be built from scratch. However, these engines are difficult to extend due to their internal complexity, and building new systems demands significant engineering effort and cost. To address this, we demonstrate GenDB, a generative query engine that shifts query processing from manually engineered systems to query processing code generation driven by Large Language Models (LLMs). An early prototype of GenDB uses LLM agents to generate instance-optimized query execution code tailored to specific data, workloads, and hardware resources. This prototype suits offline code generation for repetitive, templated queries, since the upfront generation cost amortizes over many executions and correctness can be ensured through extensive fuzz testing and manual inspection. For ad-hoc queries, GenDB can work with a traditional DBMS in a hybrid architecture: the DBMS handles one-off queries, while GenDB speeds up frequent SQL templates. Our demonstration allows users to (1) visually and interactively explore how GenDB analyzes workloads, profiles hardware resources and underlying data, produces query plans, generates code based on them, and finally uses an optimizer to iteratively achieve a correct and efficient implementation; (2) use visual inspection and analysis to gain qualitative insights into why GenDB produces code that achieves significantly better performance than state-of-the-art query engines on two benchmarks: TPC-H and a newly constructed benchmark designed to reduce potential data leakage from LLM training data; and (3) upload their own data and queries to explore GenDB with different LLMs and query patterns.
发表机构
- Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。