AI 中文总结
研究跨语言和引擎的数据组合问题,核心方法是将数据契约视为类型,通过新SDK让用户用模式对象注释表,Bauplan在执行生命周期解释注释,主要贡献是解决生产故障及实现跨语言数据流推理。
AI 中文摘要
可组合数据系统有望让开发者在不牺牲连贯用户体验的情况下组合语言、引擎和目录。但实际上,管道节点边界的规范仍很薄弱:转换通过模式交换表,这些模式检查通常滞后,跨语言执行不均衡,且与业务用户关心的语义脱节。基于在Bauplan中运行数百万个作业一年多的经验,我们分享新SDK背后的设计原则,它将数据契约视为可组合多语言数据湖的类型。用户(无论是人还是代理)用编码列类型、约束、文档和沿袭的模式对象注释输入和输出表;Bauplan然后在执行生命周期的不同点解释这些注释。我们展示了这种设计如何解决常见生产故障,以及“一切皆代码”理念如何实现对跨语言和引擎的数据流进行确定性和非确定性推理。
英文摘要
Composable data systems promise to let developers combine languages, engines, and catalogs without sacrificing a coherent user experience. In practice, however, pipeline-node boundaries remain weakly specified: transformations exchange tables through schemas that are often checked late, enforced unevenly across languages, and disconnected from the semantics business users care about. Based on over a year of operating millions of jobs in Bauplan, we share the design principles behind our new SDK, which treats data contracts as types for a composable, multi-language lakehouse. Users, whether humans or agents, annotate input and output tables with schema objects that encode column types, constraints, documentation, and lineage; Bauplan then interprets these annotations at different points in the execution lifecycle. We show how this design addresses common production failures, and how an ''everything-as-code'' philosophy enables both deterministic and non-deterministic reasoning over data flows across languages and engines.
CommentsPre-print of paper accepted at CDMS @ VLDB 2026 (Boston)