arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效可证明私密分类的表格基础模型

Efficient Provably Private Classification with a Tabular Foundation Model

Talal Alrawajfeh, Cristiana Diaconu, Ossi Räisä, Sebastian Rodriguez Beltran, Yuan He, John Bronskill, Richard E. Turner, Antti Honkela

arXiv 2610.10068首次发表:更新:

发表机构

University of Helsinki; University of Cambridge; CISPA Helmholtz Center for Information Security(赫尔辛基大学; 剑桥大学; CISPA亥姆霍兹信息安全中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PrivTab是一个内嵌隐私机制的表格基础模型,通过上下文学习生成可证明私密的摘要,在强隐私下优于现有基线,泄漏可忽略,拟合时间减少万倍。

AI 中文摘要

表格数据支撑着医学、金融、政府和科学领域的预测与决策,但往往包含敏感的个人层面信息,因此需要在保护隐私的同时实现准确预测。传统的私有学习提供了正式的隐私保证,但需要针对数据集进行缓慢的特定优化,在强隐私保护下遭受显著的效用损失,且通常难以正确应用。表格基础模型能快速适应新数据集,但现有模型缺乏正式的隐私保证,并且极易受到成员推断攻击,限制了它们在敏感数据上的使用。在此,我们提出PrivTab,一个易于使用的用于差分隐私分类的表格基础模型,其架构内嵌隐私机制。通过在模拟数据集上预训练,PrivTab利用上下文学习将敏感行转换为紧凑的、可证明私密的摘要——有效地学习如何在隐私下学习。PrivTab在中等至强隐私保护下优于私有线性基线和神经网络基线,表现出可忽略的成员泄漏,在强隐私下保持良好校准的预测,并将数据集拟合时间减少10,000倍,仅需一次前向传播。通过结合正式隐私、速度和易用性,PrivTab将人工智能的最新进展带到了敏感个人数据限制其采用的应用中。

英文摘要

Tabular data underpin prediction and decision-making in medicine, finance, government and science, but often contain sensitive individual-level information, creating a need for accurate prediction while preserving privacy. Traditional private learning provides formal privacy guarantees, but requires slow dataset-specific optimisation, suffers substantial utility loss under strong privacy, and is often difficult to apply correctly. Tabular foundation models adapt rapidly to new datasets, but existing models lack formal privacy guarantees, and are highly vulnerable to membership-inference attacks, limiting their use on sensitive data. Here we introduce PrivTab, an easy to use tabular foundation model for differentially private classification that embeds a privacy mechanism within its architecture. Pretrained on simulated datasets, PrivTab uses in-context learning to transform sensitive rows into compact, provably private summaries---effectively learning how to learn under privacy. PrivTab outperforms private linear and neural-network baselines under moderate-to-strong privacy, shows negligible membership leakage, maintains well-calibrated predictions under strong privacy, and reduces dataset fitting time by 10,000 times, requiring only a single forward pass. By combining formal privacy, speed, and easy of use, PrivTab brings recent advances in AI to applications where sensitive individual-level data have limited their adoption.

Comments74 pages, 18 figures; includes supplementary information

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑