arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PLSQLBench:面向可执行过程式数据库编程的大语言模型系统基准测试

PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming

Marianne Menglin Liu, Leonid Boytsov, Daniel W. Peterson, Pramuditha Perera, Rongguang Wang, Sai Ashish Somayajula, Syed Hamza Rafique, Rohit Saini, Shubham Pathak, Sujeeth Bharadwaj, Tao Sheng, Graham Horwood, Fahad Shah, Ankan Bansal, Sujith Ravi, Dan Roth

arXiv 2608.15931首次发表:更新:

发表机构

Oracle AI, Turing Enterprise Inc(Oracle AI 图灵企业公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出首个过程式数据库编程基准PLSQLBench,含2865个任务实例,经8个LLM实验发现其在多方面存在困难,工具增强智能体性能有提升但仍有差距,凸显了传统基准未覆盖的能力。

AI 中文摘要

我们提出了PLSQLBench,据我们所知,这是首个用于评估大语言模型(LLM)能否编写可执行PL/SQL程序的基准测试,其正确性通过基于执行的测试进行衡量。现有LLM评估大多针对通用代码生成或声明式文本转SQL,而过程式数据库编程的研究不足。PLSQLBench包含2865个实例:2594个单轮任务和271个多轮对话,共978轮。该基准结合了基于企业级Spider 2数据库的复杂模式任务、从Spider衍生的简单模式任务,以及从MBPP衍生的过程式问题,涵盖不同程度的数据库基础和过程复杂度。对8个LLM的实验显示,它们在模式基础、PL/SQL方言保真度、过程控制流、异常处理和跨轮一致性方面存在反复出现的困难。工具增强型LLM智能体在多项基于模式的评估中提升了性能,但仍存在显著差距。这些结果凸显了传统代码生成或文本转SQL基准未直接评估的过程式数据库编程能力。我们的代码可在此URL获取。

英文摘要

We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based tests. Existing LLM evaluations largely target general-purpose code generation or declarative text-to-SQL, leaving procedural database programming underexplored. PLSQLBench contains 2,865 instances: 2,594 single-turn tasks and 271 multi-turn conversations spanning 978 turns. The benchmark combines complex schema-grounded tasks over enterprise-style Spider 2 databases, simpler schema-grounded tasks derived from Spider, and MBPP-derived procedural problems, covering varying levels of database grounding and procedural complexity. Experiments with eight LLMs reveal recurring difficulties in schema grounding, PL/SQL dialect fidelity, procedural control flow, exception handling, and cross-turn consistency. Tool-augmented LLM agents improve performance on several schema-grounded evaluations, although substantial gaps remain. These results highlight procedural database programming capabilities not directly assessed by conventional code generation or text-to-SQL benchmarks. Our code is available at https://github.com/oracle-samples/plsqlbench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑