AutoSQL:从大规模代码仓库中的命令式ORM代码提取SQL模板
AutoSQL: Extracting SQL Templates from Imperative ORM Code in Large-Scale Repositories
查看机构详情
- Sun Yat-sen University(中山大学)
- Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
AutoSQL是从大规模Go仓库的命令式ORM代码提取SQL模板的系统,采用代码索引、混合上下文检索策略,在基准测试中召回率优于现有方法。
中文摘要 AI 辅助
次优SQL查询会显著降低云系统的性能,因此需要在部署前提取并审计SQL语句。然而,Go ORM框架通过分散的方法调用序列命令式地构建SQL,难以静态恢复生成的SQL模板。本文提出AutoSQL系统,用于从Go ORM代码重构SQL模板。AutoSQL构建代码索引(Code Index),这是一个有向图,将函数、类型和全局变量之间的结构依赖关系捕获为可导航边;接着从ORM调用点向上追踪调用链,识别出与数据库交互的函数作为入口点。对于每个入口点,大语言模型(LLM)智能体遍历代码索引以收集影响SQL生成的代码片段,当图无法解决检索目标时切换为基于模式的搜索,该策略称为混合上下文检索(Hybrid Context Retrieval)。收集到足够上下文后,智能体合成SQL模板。在包含5个大规模Go仓库的579个带测试覆盖的入口点和1186条运行时追踪的SQL语句的基准测试中,AutoSQL的召回率达到68.04%至72.18%,比静态可达性基线高出11.80%至15.94%,比现有方法高出8.52%至21.50%。
英文摘要
Suboptimal SQL queries can significantly degrade the performance of cloud systems, motivating the extraction and auditing of SQL statements before deployment. However, Go ORM frameworks construct SQL imperatively through scattered method-call sequences, making it difficult to statically recover the resulting SQL templates. We present AutoSQL, a system that reconstructs SQL templates from Go ORM code. AutoSQL constructs a Code Index, a directed graph that captures structural dependencies between functions, types, and global variables as navigable edges. It then traces upstream call chains from ORM invocation sites to identify database-interacting functions as entry points. For each entry point, an LLM agent traverses the Code Index to collect code slices that influence SQL generation, switching to pattern-based search when the graph cannot resolve a retrieval goal. We call this strategy Hybrid Context Retrieval. Once sufficient context is collected, the agent synthesizes SQL templates. Evaluation on a benchmark of 579 test-covered entry points and 1,186 runtime-traced SQL statements from five large-scale Go repositories shows that AutoSQL achieves 68.04% to 72.18% recall, exceeding the static reachability baseline by 11.80% to 15.94% and outperforming existing methods by 8.52% to 21.50%.