arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.06564cs.CL

CSR-RAG:面向企业级文本到SQL的高效检索系统

CSR-RAG: An Efficient Retrieval System for Text-to-SQL on the Enterprise Scale

  • Technical University of Munich(慕尼黑技术大学)
  • Nokia Bell Labs(诺基亚贝尔实验室)

机构由 AI 辅助整理,请以论文原文为准。

Rajpreet Singh, Novak Boškov, Lawrence Drabeck, Aditya Gudal, Manzoor A. Khan

更新

AI总结:

CSR-RAG通过结合上下文、结构和关系检索,高效支持企业级文本到SQL任务,实现高精度与低延迟的检索性能。

AI中文摘要:

自然语言到SQL翻译(文本到SQL)是长期存在的问题,最近因大型语言模型(LLMs)的进步而受益。尽管大多数学术文本到SQL基准要求将模式描述作为自然语言输入的一部分,但企业级应用通常需要在生成SQL查询之前进行表检索。为了解决这一需求,我们提出了一种新颖的混合检索增强生成(RAG)系统,由上下文、结构和关系检索(CSR-RAG)组成,以实现对企业级数据库的计算高效且足够准确的检索。通过广泛的enterprise基准测试,我们证明CSR-RAG在商品数据中心硬件上实现了高达40%的精度和超过80%的召回率,同时平均查询生成延迟仅为30毫秒,这使其适合现代基于LLM的企业级系统。

英文摘要:

Natural language to SQL translation (Text-to-SQL) is one of the long-standing problems that has recently benefited from advances in Large Language Models (LLMs). While most academic Text-to-SQL benchmarks request schema description as a part of natural language input, enterprise-scale applications often require table retrieval before SQL query generation. To address this need, we propose a novel hybrid Retrieval Augmented Generation (RAG) system consisting of contextual, structural, and relational retrieval (CSR-RAG) to achieve computationally efficient yet sufficiently accurate retrieval for enterprise-scale databases. Through extensive enterprise benchmarks, we demonstrate that CSR-RAG achieves up to 40% precision and over 80% recall while incurring a negligible average query generation latency of only 30ms on commodity data center hardware, which makes it appropriate for modern LLM-based enterprise-scale systems.

↑