Reflect-SQL:一种基于自反思的Text-to-SQL框架
Reflect-SQL: A Self-Reflection Based Framework for Text-to-SQL
- International Institute of Information Technology Hyderabad(国际信息技术研究所海得拉巴分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Reflect-SQL是一种基于多阶段自反思的Text-to-SQL框架,通过LLM作为评判者的反馈循环优化结果,在BIRD基准上实现72.03%执行准确率,显著优于现有方法,提升了企业数据访问的可靠性。
AI中文摘要:
通过自然语言实现数据访问的民主化是现代企业的关键目标,但Text-to-SQL的实际应用却受到现实复杂性的严重阻碍:1. 晦涩且庞大的数据库模式;2. 由于模式的结构化设置和用户查询模糊,导致无法有效检索相关的表和列;3. 由于缺乏可靠的验证和修正机制,生成的SQL存在语法或逻辑缺陷。为应对这些系统性挑战,我们提出了Reflect-SQL,一种用于Text-to-SQL的新颖框架,该框架基于多阶段自反思方法,利用知识库理解晦涩的模式,建立有效检索流程,并生成符合语法和语义的SQL。我们的系统不采用单次生成,而是在相互关联的反馈循环中使用“大语言模型作为评判者”的评分机制,在每个阶段迭代优化结果:反馈驱动的检索循环会优化用户的自然语言查询,合成循环会验证并修正SQL,最终蕴含循环会优化端到端流程并持续丰富知识库。通过整合这些反思层级,Reflect-SQL弥合了用户意图与复杂数据之间的关键差距。在具有挑战性的BIRD基准上,我们的框架实现了72.03%的执行准确率,显著优于最先进的基线模型,证明其在企业应用可靠性上取得了重大突破。
英文摘要:
Democratizing data access through natural language is a crucial goal for modern enterprises, but the practical adoption of Text-to-SQL is critically hindered by real-world complexities: 1. Obscure and large database schemas, 2. Ineffective retrieval of relevant tables and columns due to structured setting of schemas and vague user query, 3. Generation of syntactically or logically flawed SQL due to a lack of robust validation and correction mechanism. To address these systemic challenges, we introduce Reflect-SQL, a novel framework for Text to SQL, grounded in multi-stage self-reflection approach to develop understanding of obscure schema using a knowledge base, setup a process for effective retrieval and system to generate syntactically/semantically SQL. Instead of a single-pass attempt, our system employs an LLM-as-a-judge driven scoring mechanism within interconnected feedback loops to iteratively refine the results at every stage. A feedback-driven retrieval loop refines the user's natural language query, while a synthesis loop validates and corrects the SQL and finally, an entailment loop optimizes the end-to-end process and continuously enriches the knowledge base. By integrating these layers of reflection, Reflect-SQL bridges the critical gap between user intent and complex data. On the challenging BIRD benchmark, our framework achieves an execution accuracy of 72.03%, significantly outperforming state-of-the-art baselines, demonstrating a major leap in reliability for enterprise applications.