发表机构
German Cancer Research Center (DKFZ); Medical Faculty Heidelberg, Heidelberg University; Bilkent University; Pattern Analysis and Learning Group, Department of Radiation Oncology, Heidelberg University Hospital(德国癌症研究中心; 海德堡大学医学院; 比尔肯特大学; 海德堡大学医院放射肿瘤学系模式分析与学习组)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究放射学报告联邦学习中的隐私泄露,通过固定模型架构比较三种分词器的隐私风险,量化梯度文本重建,发现分词器设计影响泄露程度,虽RadBERT重建保真度最高但均不能防泄露,安全聚合等保障措施或为满足相关要求所必需。
AI 中文摘要
联邦学习(FL)能够在不共享原始数据的情况下进行多机构临床文本训练,但梯度反转可以从共享的模型更新中重建敏感信息。放射学报告中这种泄露的程度以及分词器设计的作用仍不明确。我们量化了联邦学习中基于梯度的文本重建,并在固定模型架构的情况下比较了三种分词器的隐私风险。六个联邦学习客户端使用GPT-2、RadBERT和LLaMA-2分词器,在公共放射学语料库上训练GPT-2风格的变压器(序列长度32),批量大小分别为64、128和256。假设存在一个活跃的恶意服务器,在分发前修改共享架构,我们应用分析梯度反转并在五次运行中测量重建保真度。跨分词器的精确句子重建范围为31%至44%。在出院数据集上,批量大小为64时,准确率分别为42.1%(GPT-2)、42.3%(RadBERT)和39.4%(LLaMA-2),批量大小为256时降至37.3%、37.2%和34.3%。S-BLEU随着批量大小的增加而下降。RadBERT产生了最高的重建保真度,恢复了最多的临床术语,但没有一个分词器能防止泄露。因此,即使在更大的批量大小和使用特定领域分词器的情况下,报告文本的很大一部分仍可从联邦学习梯度中恢复。分词器设计影响泄露严重程度,是一个与隐私相关的决策,而不仅仅是实用性决策;安全聚合和差分隐私等保障措施可能是满足放射学自然语言处理中联邦学习的HIPAA和GDPR要求所必需的。
英文摘要
Federated learning (FL) enables multi-institutional training on clinical text without sharing raw data, but gradient inversion can reconstruct sensitive information from shared model updates. The extent of this leakage for radiology reports, and the role of tokenizer design, remains unclear. We quantify gradient-based text reconstruction in FL and compare privacy risk across three tokenizers with the model architecture held fixed. Six FL clients trained a GPT-2-style transformer (sequence length 32) on public radiology corpora (368,751 diagnostic reports, 98,206 discharge summaries, 1,500 MIMIC-CXR free-text reports) using the GPT-2, RadBERT, and LLaMA-2 tokenizers at batch sizes of 64, 128, and 256. Assuming an active malicious server that modifies the shared architecture before distribution, we applied analytic gradient inversion and measured reconstruction fidelity over five runs. Exact sentence reconstruction ranged from 31% to 44% across tokenizers (30.6-43.5% across the 27 tokenizer x dataset x batch-size cells). At batch size 64 on the Discharge dataset, accuracy was 42.1% (GPT-2), 42.3% (RadBERT), and 39.4% (LLaMA-2), decreasing to 37.3%, 37.2%, and 34.3% at batch size 256. S-BLEU declined as batch size grew (GPT-2: 0.44 to 0.33; RadBERT: 0.48 to 0.35). RadBERT yielded the highest reconstruction fidelity and recovered the most clinical terms (18.1% of a 1,440-term reference vocabulary, vs 12.5% for GPT-2 and 9.4% for LLaMA-2), yet no tokenizer prevented leakage. Substantial portions of report text are therefore recoverable from FL gradients even at larger batch sizes and with domain-specific tokenizers. Tokenizer design influences leakage severity and is a privacy-relevant decision, not only a utility one; safeguards such as secure aggregation and differential privacy are likely necessary to meet HIPAA and GDPR requirements for FL in radiology NLP.