发表机构
Oracle AI, Turing Enterprise Inc(Oracle AI 图灵企业公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出首个过程式数据库编程基准PLSQLBench,含2865个任务实例,经8个LLM实验发现其在多方面存在困难,工具增强智能体性能有提升但仍有差距,凸显了传统基准未覆盖的能力。
AI 中文摘要
我们提出了PLSQLBench,据我们所知,这是首个用于评估大语言模型(LLM)能否编写可执行PL/SQL程序的基准测试,其正确性通过基于执行的测试进行衡量。现有LLM评估大多针对通用代码生成或声明式文本转SQL,而过程式数据库编程的研究不足。PLSQLBench包含2865个实例:2594个单轮任务和271个多轮对话,共978轮。该基准结合了基于企业级Spider 2数据库的复杂模式任务、从Spider衍生的简单模式任务,以及从MBPP衍生的过程式问题,涵盖不同程度的数据库基础和过程复杂度。对8个LLM的实验显示,它们在模式基础、PL/SQL方言保真度、过程控制流、异常处理和跨轮一致性方面存在反复出现的困难。工具增强型LLM智能体在多项基于模式的评估中提升了性能,但仍存在显著差距。这些结果凸显了传统代码生成或文本转SQL基准未直接评估的过程式数据库编程能力。我们的代码可在此URL获取。
英文摘要
We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based tests. Existing LLM evaluations largely target general-purpose code generation or declarative text-to-SQL, leaving procedural database programming underexplored. PLSQLBench contains 2,865 instances: 2,594 single-turn tasks and 271 multi-turn conversations spanning 978 turns. The benchmark combines complex schema-grounded tasks over enterprise-style Spider 2 databases, simpler schema-grounded tasks derived from Spider, and MBPP-derived procedural problems, covering varying levels of database grounding and procedural complexity. Experiments with eight LLMs reveal recurring difficulties in schema grounding, PL/SQL dialect fidelity, procedural control flow, exception handling, and cross-turn consistency. Tool-augmented LLM agents improve performance on several schema-grounded evaluations, although substantial gaps remain. These results highlight procedural database programming capabilities not directly assessed by conventional code generation or text-to-SQL benchmarks. Our code is available at https://github.com/oracle-samples/plsqlbench.