arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

国家规模神经影像存储库中的身份重复审计

Identity-Duplication Auditing in National-Scale Neuroimaging Repositories

Jiheng Li, Michael E. Kim, Trent M. Schwartz, Yuhan Cui, Gaurav Rudravaram, Derek B. Archer, Timothy J. Hohman, Lori L. Beason-Held, Victoria L. Morgan, Dario J. Englot, Angela L. Jefferson, for the Alzheimer's Disease Neuroimaging Initiative, for the BIOCARD Study team, for the Health, Aging Brain Study, :, Health Disparities, Study Team, Lianrui Zuo, Guray Erus, Christos Davatzikos, Bennett A. Landman

arXiv 2610.09614首次发表:更新:

发表机构

Vanderbilt University; University of Pennsylvania; Vanderbilt Health; National Institute on Aging, National Institutes of Health(范德堡大学; 宾夕法尼亚大学; 范德堡医疗中心; 美国国立卫生研究院国家老龄化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出HAPPEN,一种人机协同流水线,用于在国家规模T1加权脑MRI存储库中审计身份重复,结合SHA-256指纹与监督对比检索,在95,129次扫描中识别出重复组并实现全部遗传参考配对的恢复。

AI 中文摘要

国家规模的磁共振成像(MRI)存储库日益整合来自不同研究和机构的数据。然而,仅在单个数据集内有效的受试者标识符在聚合后不再保证全局唯一,这使得同一受试者可能被分配多个标识符,我们将其定义为身份重复。这种重复会在训练数据和测试数据之间造成泄漏,并夸大下游生物医学研究中的表观性能。现有方法未提供端到端的、基于图像的工作流来在存储库规模上审计此问题。在本工作中,我们提出了HAPPEN,一种用于审计T1加权脑MRI存储库中身份重复的人机协同流水线。它结合了SHA-256指纹识别用于精确重复检测,以及监督对比检索用于可能来自同一人的非相同扫描。检索到的配对在本地托管的界面中作为候选进行审查,而非自动分类为重复。我们在一个包含95,129次扫描的聚合存储库中部署了该工作流,并使用54个遗传参考配对评估了端到端恢复率。可迁移性通过在独立机构的22,386次扫描上本地部署相同工作流进行评估,无需模型重新训练或图像传输。在研究存储库中的部署识别出1,316个精确重复扫描组和1,275个审查者支持的近似重复受试者组。在这些组中,分别有56%和82%跨越了数据集边界。所有54个遗传参考配对均被恢复。外部团队使用本地选择的操作阈值和审查标准独立完成了整个工作流。

英文摘要

National-scale magnetic resonance imaging (MRI) repositories increasingly integrate data from different studies and institutions. However, subject identifiers that are valid only within individual datasets are no longer guaranteed to remain globally unique after aggregation, making it possible for the same subject to be assigned multiple identifiers, which we define as identity duplication. Such duplication can create leakage between training and test data and inflate apparent performance in downstream biomedical studies. Existing methods do not provide an end-to-end, image-based workflow for auditing this problem at repository scale. In this work, we present HAPPEN, a human-in-the-loop pipeline for auditing identity duplication in T1-weighted brain MRI repositories. It combines SHA-256 fingerprinting for exact-duplicate detection with supervised contrastive retrieval of non-identical scans that may originate from the same person. Retrieved pairs are reviewed as candidates in a locally hosted interface rather than automatically classified as duplicates. We deployed the workflow in a 95,129-scan aggregated repository and assessed end-to-end recovery using 54 genetic-reference pairs. Transferability was assessed by locally deploying the same workflow on 22,386 scans at an independent institution without model retraining or image transfer. Deployment in the study repository identified 1,316 exact-duplicate scan groups and 1,275 reviewer-supported near-duplicate subject groups. Of these groups, 56% and 82%, respectively, crossed dataset boundaries. All 54 genetic-reference pairs were recovered. The external team independently completed the full workflow using a locally selected operating threshold and review standard.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑