发表机构
TU Darmstadt; University of Edinburgh; EPFL; RelationalAI(达姆施塔特工业大学; 爱丁堡大学; 洛桑联邦理工学院; 关系人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出将线性 Datalog 程序编译为递归 SQL 的框架,通过中间语言 Midlog 和函数依赖恢复,使现有数据库引擎高效执行程序分析,在 Umbra 上比 Soufflé 快一个数量级。
AI 中文摘要
Datalog 是一种声明式查询语言,已被证明在表达静态程序分析方面非常有效。尽管 Datalog 深深植根于数据库理论,但最近的大多数进展主要来自编程语言和编译器社区,例如 Soufflé 系统。相比之下,现代关系引擎在优化递归 SQL 方面取得了显著进展。本文重新审视了 Datalog 与关系数据库之间的联系,提倡将递归 SQL 作为 Datalog 评估的后端。我们提出了一个编译框架,将 Datalog 程序(尤其是线性 Datalog 片段中的程序)翻译为等价的递归 SQL 查询。为了弥合 Datalog 与 SQL 之间的差距,编译器将每个程序都通过一个名为 Midlog 的中间语言进行路由。此外,编译器从程序中恢复函数依赖,并将其暴露为模式键,从而释放引擎的标准查询优化能力。这种方法使得现有数据库引擎能够执行广泛的程序分析,在 Umbra 后端上性能比 Soufflé 引擎高出一个数量级。Umbra 在 8 线程时实现了 5.46 倍的几何平均加速比,而 DuckDB 在单线程下与 Soufflé 相当,但在 8 线程时较慢(几何平均加速比为 0.68 倍)。此外,生成的 SQL 是可移植的;它可以在七个数据库系统上运行,无需修改引擎。我们的结果突显了关系引擎全面支持 Datalog 以进行大规模程序分析所需的条件。
英文摘要
Datalog is a declarative query language that has proven highly effective for expressing static program analyses. Although Datalog has deep roots in database theory, most recent advances have largely emerged from the programming languages and compiler communities, with systems such as Soufflé. In contrast, modern relational engines have made significant progress in optimizing recursive SQL. This paper revisits the connection between Datalog and relational databases, advocating recursive SQL as a backend for Datalog evaluation. We present a compilation framework that translates Datalog programs, particularly those in the Linear Datalog fragment, into equivalent recursive SQL queries. To bridge the gap between Datalog and SQL, the compiler routes every program through an intermediate language called Midlog. The compiler additionally recovers functional dependencies from the program and exposes them as schema keys, unlocking the engine's standard query optimizations. This approach enables existing database engines to execute a broad class of program analyses, outperforming the Soufflé engine by up to an order of magnitude on the Umbra backend. Umbra achieves a geometric-mean speedup of 5.46$\times$ at 8 threads, whereas DuckDB is competitive with Soufflé single-threaded and is slower at 8 threads (geometric-mean speedup of 0.68$\times$). Furthermore, the generated SQL is portable; it runs on seven database systems without any engine modification. Our results highlight what the relational engines require to fully support Datalog for large-scale program analysis.