arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

H2:一种用于医疗数据协调的双混合语义数据湖架构,带有人类在环验证、大语言模型驱动的元数据标注系统

H2: A Dual Hybrid Semantic Data Lake Architecture for Medical Data Harmonization with Human-In-the-Loop verified, LLM Driven Metadata Annotation System

Ioannis N. Tzortzis, Georgia Kapetadimitri, Agapi Davradou, Nefeli Kousta, Nikolaos Bakalos, Ioannis Rallis, Dimitrios Kalogeras, Nikolaos Doulamis, Anastasios Doulamis

arXiv 2608.08056首次发表:更新:

发表机构

Institute of Communication and Computer Systems (ICCS); University of Macedonia (UOM)(通信与计算机系统研究所(ICCS); 马其顿大学(UOM))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出带人类在环验证、LLM驱动元数据标注系统的双混合语义数据湖架构,解决医疗数据异质性与数据沼泽问题,支撑医疗数据协调及机器学习技术应用。

AI 中文摘要

医疗数据本质上在多个层面表现出高度异质性,包括:(a) 图像、文本和时间序列等不同模态;(b) 各机构引入的多样化表格模式;(c) 医疗保健专业人员提供的完全非结构化文本信息。数据湖常被用于医疗数据存储,以将所有异构多样化数据整合到单个中央位置,可“原样”保存数据,无需像数据仓库那样施加模式。尽管数据湖具有灵活性,但却因“数据沼泽”故障而声名狼藉。因此,在不损害完整性或灵活性的前提下,通过元数据提供可靠的数据协调机制是一项真正的挑战。为此,知识图谱因能以动态方式描绘关系且无需严格的写入时模式方法而受到关注。此外,另一项严格的任务依赖于数据的互操作性:对这种多样化数据应用合适的机器学习技术并非易事,因为领域专家必须确定某一方法对特定数据类型或数据集的有效性。元数据标注可通过标记适用操作提供帮助,但这需要人工干预,更不用说大量现有数据集缺乏此类信息。为应对这两项挑战,本文提出一种语义数据湖架构,该架构可促进数据协调,并结合非标注元数据集合的生成式标注过程(即大语言模型),以支持有意义的机器学习技术的应用。在此方法基础上,我们构建更高层次的知识,基于数据的性质确定数据对适用机器学习操作的适用性……

英文摘要

Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, text and time series, (b) diverse tabular schemata introduced by institutions and (c) completely unstructured textual information data provided by healthcare professionals. Data lakes are often used in medical data storage to consolidate all heterogeneous diverse data in a single, central location, where it can be saved "as is", without the need to impose a schema like a data warehouse does. Despite their flexibility, though, data lakes are notorious for the "data swamp" failure. Thus, providing a reliable data harmonization mechanism through metadata, without compromising integrity or flexibility, is a real challenge. To this end, knowledge graphs have attracted attention since they provide a dynamic way to depict relationships without a rigid schema-on-write approach. Additionally, another rigorous task relies on the interoperability of data: application of appropriate ML techniques on such a diverse nature of data is not an easy task, as a domain expert must decide the efficacy of a method to a specific data type or dataset. Metadata annotation can aid by tagging applicable operations, however this requires manual intervention, not to mention the plethora of existing datasets which lack such information. To tackle both challenges, in this paper, we propose a semantic data lake architecture that promotes data harmonization and incorporates a generative annotation process (i.e. LLMs) of non-labeled metadata collections to support the application of meaningful ML techniques. Building on top of this approach, we create a higher level of knowledge, identifying suitability of data with respect to applicable ML operations based on their data nature...

CommentsThis work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑