arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用大语言模型发现数据湖中的关系:一个工业案例

Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case

Ahlame Diouan, Eric Ferey, Sabine Loudcher, Jérôme Darmont

arXiv 2608.26750首次发表:更新:

AI 中文总结

针对数据湖元数据对列关系发现信息不足的问题,提出两阶段方法ColRel,结合元数据、数据及业务字典构建列嵌入,在公共基准和工业ERP数据集上验证其在语义相关弱信号场景中效果显著。

AI 中文摘要

数据湖依赖元数据保持可用性,但元数据往往有限或对列关系发现的信息性较弱,尤其在源自ERP的数据集(其模式标签为编码或缩写形式)中。我们提出ColRel这一两阶段方法,在数据摄入时从元数据和可用数据构建列嵌入。在编码模式等困难场景中,业务字典可帮助更好地解释列名,并支持生成用于第二阶段的简短自然语言描述。在公共基准和工业ERP数据集上的实验表明,ColRel在语义相关、弱信号场景中效果尤为显著。

英文摘要

Data lakes rely on metadata to remain usable, yet this meta data is often limited or weakly informative for column relationship discovery, especially in ERP-derived datasets with coded or abbreviated schema labels. We propose ColRel, a two-stage method that builds column embeddings from metadata and data available at ingestion time. In difficult cases, such as coded schemata, business dictionaries help better interpret column names and support the generation of short natural-language descriptions used in the second stage. Experiments on public benchmarks and an industrial ERP dataset show that ColRel is particularly effective in semantically related, weak-signal settings.

Journal ref28th International Conference on Big Data Analytics and Knowledge Discovery (DaWaK 2026), Aug 2026, Graz, Austria. pp.116-130

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑