AI 中文总结
针对联邦环境中数据共享管道激增的问题,提出以复用为指导原则,定义复用内涵、识别复用时机,给出面向复用的设计方法与参考架构,初步评估显示复用可减少冗余、提升可管理性。
AI 中文摘要
数据网格架构通过域所有的数据产品实现去中心化数据共享,但在联邦环境中支持不同的消费者往往需要定制化的数据共享管道。随着消费者数量的增长,这会导致管道激增,增加设计与维护的复杂度。我们观察到这类管道常存在大量结构重叠,因此提出复用应作为解决该挑战的指导原则:将数据共享管道中的复用定义为跨管道系统性使用现有数据资产与转换逻辑,并在设计时与运行时识别复用机会。我们分析了相关挑战,概述了一种面向复用的设计方法,辅以参考架构。初步评估显示,复用有望减少冗余、提升可管理性,为实现更具可扩展性与可持续性的联邦数据共享提供了途径。
英文摘要
Data mesh architectures enable decentralized data sharing through domain-owned data products, but supporting diverse consumers in federated settings often requires customized data-sharing pipelines. As the number of consumers grows, this leads to a proliferation of pipelines, increasing design and maintenance complexity. We observe that such pipelines frequently exhibit substantial structural overlap. In this paper, we argue that reuse should serve as a guiding principle to address this challenge. We define reuse in data-sharing pipelines as the systematic use of existing data assets and transformation logic across pipelines, and identify reuse opportunities at both design time and runtime. We analyze the associated challenges and outline a reuse-oriented design approach, supported by a reference architecture. A preliminary evaluation demonstrates the potential of reuse to reduce redundancy and improve manageability, providing a pathway toward more scalable and sustainable federated data sharing.