DAGSmith:面向dbt风格SQL流水线的依赖感知重写
DAGSmith: Dependency-Aware Rewriting for dbt-Style SQL Pipelines
浏览论文内容
中文总结 AI 辅助
DAGSmith是首个面向dbt风格SQL流水线DAG的依赖感知源到源重写系统,在Tuva dbt项目上可缩短运行时间、降低计算成本,效果优于现有单查询重写方法。
中文摘要 AI 辅助
现代分析工作日益以循环SQL流水线的形式组织,而非孤立的SQL语句。近年来大受欢迎的dbt等工具允许团队将每个转换编写为SQL,并明确转换间的依赖关系,从而生成包含成百上千个相互依赖SQL模型的有向无环图(DAG)。传统查询优化器和源到源查询重写器一次仅处理一个查询,而物化视图选择和多查询优化仅针对更狭窄的复用形式,它们未利用明确依赖所暴露的流水线级信息:中间结果的使用方式、每个计算依赖哪些下游输出、昂贵操作相对于数据缩减的位置、哪些结果值得持久化,以及刷新计划如何关联输入变化与输出需求。我们推出DAGSmith,据我们所知,这是首个面向SQL流水线DAG的整体式依赖感知源到源重写系统。DAGSmith将明确依赖视为优化信号,它结合上游输入、下游使用者及在流水线DAG中的位置分析每个转换,利用大语言模型(LLM)提出流水线级重构方案,分离SQL生成与等价性检查以拒绝不安全重写,通过学习到的代价模型调整持久化选择,并选取全局兼容、无冲突的重写集合。这可实现依赖边简化、非局部语义复用、下游感知剪枝、流水线感知工作放置、重写-物化协同优化及频率感知优化。在开源Tuva dbt项目上,DAGSmith将运行时间缩短42.6%,仓库计算成本降低67.7%,分别比最先进的单查询重写方法提升98.1%和348.3%。
英文摘要
Modern analytics is increasingly organized as recurring SQL pipelines rather than isolated SQL statements. Tools such as dbt, which have gained extreme popularity in recent years, allow teams to write each transformation as SQL and make dependencies between transformations explicit, producing directed acyclic graphs (DAGs) with hundreds or thousands of interdependent SQL models. Traditional query optimizers and source-to-source query rewriters operate on one query at a time, while materialized-view selection and multi-query optimization address narrower forms of reuse. They do not exploit the pipeline-level information exposed by explicit dependencies: how intermediate results are consumed, which downstream outputs depend on each computation, where expensive work sits relative to data reduction, which results are worth persisting, and how refresh schedules relate to input change and output demand. We introduce DAGSmith, to the best of our knowledge the first holistic dependency-aware source-to-source rewriting system for SQL pipeline DAGs. DAGSmith treats explicit dependencies as optimization signals. It analyzes each transformation with its upstream inputs, downstream consumers, and position in the pipeline DAG, uses an LLM to propose pipeline-level refactorings, separates SQL generation and equivalence checking to reject unsafe rewrites, retunes persistence choices with a learned cost model, and selects a globally compatible, conflict-free set of rewrites. This enables dependency-edge simplification, non-local semantic reuse, downstream-aware pruning, pipeline-aware work placement, rewrite-materialization co-optimization, and frequency-aware optimization. On the open-source Tuva dbt project, DAGSmith reduces elapsed time by 42.6% and warehouse compute cost by 67.7%, 98.1%/348.3% larger than state-of-the-art single-query rewriting.