arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

审计合成基因表达数据的隐私性:一种用于无盒成员推断的统一加权距离框架

Auditing the Privacy of Synthetic Gene Expression Data: A Unified Weighted-Distance Framework for No-Box Membership Inference

Owen Tucker, Lily Wang, Harutoshi Okumura, Ruixuan Liu, Li Xiong

arXiv 2610.04060首次发表:更新:

发表机构

Emory University(埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出统一加权距离框架审计合成基因表达数据的隐私,通过五种无盒成员推断攻击评估,发现谱残差化显著提升攻击效能,而通路先验不可迁移。

AI 中文摘要

合成基因表达数据日益被提出作为受控访问基因组存储库的隐私保护替代品,但其安全性取决于经验审计。成员推断攻击(MIA)通过测试患者的基因表达谱是否被用于训练生成模型来提供这种审计。我们报告了一项红队研究,该研究针对由ELSA Health 2026挑战赛发布的、源自癌症基因组图谱(TCGA)中批量RNA-seq谱的合成基因表达数据。我们将五种无盒攻击统一在一个加权距离框架下,其中每种变体仅在基因加权方式上有所不同:均匀加权(基线)、按方差加权、按合成与参考之间的KL散度加权、通过谱残差化去除主要主成分加权,以及按精选通路成员关系加权。针对条件变分自编码器,谱残差化在泛癌TCGA上将AUC从0.8251提升至0.8951,并将1%假阳性率下的真阳性率从0.3842提升至0.5970。生物学通路先验未能在不同目标之间转移。

英文摘要

Synthetic gene expression data is increasingly proposed as a privacy-preserving substitute for controlled-access genomic repositories, but its safety depends on empirical auditing. Membership inference attacks (MIAs) provide that audit by testing whether a patient's gene expression profile was used to train a generative model. We report a red-team study on synthetic gene expression data derived from bulk RNA-seq profiles in The Cancer Genome Atlas, released by the ELSA Health 2026 Challenge. We unify five no-box attacks under a single weighted-distance framework in which each variant differs only in how it weights genes: uniformly (baseline), by variance, by synthetic-versus-reference KL divergence, by spectral residualization removing dominant principal components, and by curated pathway membership. Against a conditional variational autoencoder, spectral residualization raises AUC from 0.8251 to 0.8951 and TPR at 1% FPR from 0.3842 to 0.5970 on pan-cancer TCGA. Biological pathway priors did not transfer across targets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑