大型语言模型能否规范化数据库?一个用于模式规范化的基准与多智能体框架
Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization
浏览论文内容
中文总结 AI 辅助
提出DNBENCH基准和MARS多智能体框架,评估并提升LLM在数据库模式规范化(1NF到BCNF)中的可靠性,MARS将证据提取、诊断和规划分离,使DNB-SCORE提升82.0%。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地被用于生成结构化输出,但当这些输出必须满足数据库级别的约束时,其可靠性仍不明确。我们通过数据库规范化来研究这一问题,这涉及对函数依赖、无损连接分解和表间约束的推理。我们引入了一个数据库规范化基准(DNBENCH),包含3,275个样本,用于评估从1NF到BCNF的LLM驱动的数据库规范化。DNBENCH采用三轴协议来衡量语义等价性、结构准确性和逻辑有效性。在单一、复杂和现实世界三个级别上,DNBENCH揭示了依赖推断、模式分解和表间约束重建中的反复失败。我们进一步提出了多智能体模式推理(MARS),它将证据提取、违规诊断和分解规划与模式生成和验证分离。MARS在DNB-SCORE上比单提示基线提高了82.0%。所有工件将在接收后发布。
英文摘要
Large Language Models (LLMs) are increasingly used to generate structured outputs, but their reliability remains unclear when those outputs must satisfy database-level constraints. We study this issue through database normalization, involving reasoning about functional dependencies, lossless join decompositions, and inter-table constraints. We introduce a Database Normalization Benchmark (DNBENCH), comprising 3,275 samples for evaluating LLM-driven database normalization from 1NF to BCNF. DNBENCH uses a three-axis protocol to measure semantic equivalence, structural accuracy, and logical validity. Across Single, Complex, and Real World levels, DNBENCH uncovers recurring failures in dependency inference, schema decomposition, and inter-table constraint reconstruction. We further propose Multi-Agent Reasoning for Schemas (MARS), which separates evidence extraction, violation diagnosis, and decomposition planning from schema generation and verification. MARS improves the DNB-SCORE by 82.0% over the single-prompt baseline. All artifacts will be released upon acceptance.
发表机构
- Kyungpook National University(庆北国立大学)
机构由 AI 辅助整理,请以论文原文为准。