arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于图方法的临床文档一致性再词汇化

Consistent Relexicalization of Clinical Documents using Graph-Based Approach

Dipankar Das, Atri Mandal, Sandeep Singh, Tushar Shandhilya

arXiv 2609.21387首次发表:更新:

发表机构

Oracle Health AI(甲骨文健康人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对临床文档再词汇化中结构、关系和时间一致性问题,提出基于图与LLM的G-RELIC方法,显著提升关系完整性和时间一致性,同时保持隐私安全。

AI 中文摘要

再词汇化是临床自然语言处理中的一项关键技术,因为它有助于在合成保留高保真、真实世界特征的数据集的同时,对敏感信息进行稳健的掩蔽。然而,在转换过程中保持结构完整性、关系连贯性和时间一致性仍然是一个重大挑战。现有方法通常依赖于独立的实体替换,这导致纵向记录中的临床不一致性。这降低了此类再词汇化数据集对下游科学分析的价值。为了解决这些局限性,我们引入了G-RELIC(基于图的上下文再词汇化与改进一致性),它结合了LLMs和图的力量。G-RELIC实现了一种基于图的映射机制,优化原始实体与替代实体之间的一一对应关系。它还引入了一种确定性时间重定位算法以保持时间一致性。在多样化的真实世界临床数据集上的实证评估验证了G-RELIC显著优于最先进的基线。G-RELIC在关系完整性方面提高了30.4个百分点(从62.1%提升至92.5%),在时间一致性方面提高了45.9个百分点(从46%提升至91.9%),且不损害公认的临床数据集隐私基准。这最大化了再词汇化数据集的分析效用,同时最小化了重新识别风险。

英文摘要

Relexicalization is a pivotal technique in clinical NLP, as it facilitates robust masking of sensitive information while synthesizing datasets that retain high-fidelity, real-world characteristics. However, preserving structural integrity, relational coherence, and temporal consistency during transformation remains a significant challenge. Existing approaches frequently rely on independent entity replacement, which results in clinical inconsistencies across longitudinal records. This reduces the value of such relexicalized datasets for downstream scientific analysis. To address these limitations, we introduce G-RELIC (Graph Based Contextual Relexicalization with Improved Consistency) which combines the power of LLMs with graphs. G-RELIC implements a graph-based mapping mechanism which optimizes for one-to-one correspondence between original and surrogate entities. It also introduces a deterministic temporal repositioning algorithm to preserve temporal consistency. Empirical evaluations on diverse, real-world clinical datasets validate that G-RELIC significantly outperforms state-of-the-art baselines. G-RELIC yields a 30.4 percentage point improvement in relational integrity (62.1% to 92.5%) and 45.9 percentage point improvement in temporal coherence (46% to 91.9%) without compromising on the recognized privacy benchmarks for clinical datasets. This maximizes the analytical utility of relexicalized datasets while minimizing re-identification risk.

CommentsAccepted for presentation at the Sixth International Conference on AI ML Systems (AIMLSystems 2026), Lake Como, Italy, October 6-9, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑