ProcArena:面向自然语言直接与交互式PL/SQL开发的大语言模型多场景基准
ProcArena: A Multi-Scenario Benchmark for LLMs on Direct and Interactive PL/SQL Development from Natural Language
浏览论文内容
中文总结 AI 辅助
本文提出ProcArena,一个覆盖直接与交互模式、含3998个可执行任务的多场景PL/SQL开发基准,评估七种大语言模型,结果显示最佳平均得分仅62.2%和57.8%,表明该任务仍具挑战性。
中文摘要 AI 辅助
大语言模型(LLMs)在将自然语言(NL)需求转换为PL/SQL程序方面展现出强大潜力,引起了数据库社区的日益关注。然而,现有的自然语言到PL/SQL(NL-to-PL/SQL)研究主要集中于从完整的自然语言需求直接生成PL/SQL。在实践中,PL/SQL开发涉及多种场景,如从零开发、代码修改、调试和优化,并且可能需要直接生成或多轮交互。然而,目前尚无全面的基准来评估多场景、直接与交互式以及多方言的NL-to-PL/SQL开发。在本文中,我们提出了ProcArena,一个基于执行的基准,涵盖直接(Direct)和交互式(Interactive)两种模式。ProcArena包含157个数据库上的3,998个可执行任务,涵盖PostgreSQL和Oracle中的九种开发子场景。我们通过迭代逻辑增强和场景特定适配器构建具有挑战性的直接任务,并通过知识整合和需求扰动推导出配对的交互式任务,同时保留可执行目标。我们进一步设计了一个受控的求解器-用户模拟器协议,允许模型澄清用户意图并检查数据库环境,而不暴露隐藏的执行反馈。评估了七种语言模型,我们发现直接和交互式模式下的最佳平均得分分别仅为62.2%和57.8%,这表明现实中的NL-to-PL/SQL开发仍然具有挑战性,尤其是在交互式设置中。
英文摘要
Large language models (LLMs) have shown strong potential for translating natural-language (NL) requirements into PL/SQL programs, attracting increasing attention from the database community. However, existing NL-to-PL/SQL efforts primarily focus on directly generating PL/SQL from complete NL requirements. In practice, PL/SQL development involves diverse scenarios, such as from-scratch development, code modification, debugging, and optimization, and may require either direct generation or multi-turn interaction. Yet, no comprehensive benchmark evaluates multi-scenario, direct and interactive, and multi-dialect NL-to-PL/SQL development. In this paper, we present ProcArena, an execution-based benchmark covering both Direct and Interactive modes. ProcArena comprises 3,998 executable tasks over 157 databases, spanning nine development subscenarios in PostgreSQL and Oracle. We construct challenging Direct tasks through Iterative Logic Enhancement and scenario-specific adapters, and derive paired Interactive tasks through Knowledge Integration and Requirement Perturbation while preserving executable targets. We further design a controlled Solver-User Simulator protocol that allows models to clarify user intent and inspect the database environment without exposing hidden execution feedback. Evaluating seven language models, we find that the best average scores are only 62.2% and 57.8% in Direct and Interactive, respectively, demonstrating that realistic NL-to-PL/SQL development remains challenging, particularly in interactive settings.
发表机构
- Tsinghua University(清华大学)
- Lenovo Group Limited(联想集团有限公司)
机构由 AI 辅助整理,请以论文原文为准。