AI 中文总结
本研究提出基于Shapley值的特征归因框架,在特征层面解决数据隐私的风险-效用权衡,可适配各类方法,实验证实其能在保留数据效用的同时降低披露风险。
AI 中文摘要
尽管个人数据的广泛访问带来诸多益处,但也引发了消费者、企业和政策制定者面临的严重隐私问题。本研究提出一种新框架,将基于Shapley值的特征归因方法应用于数据隐私领域,同时捕捉数据隐私的两个关键维度:披露风险和数据效用。该框架通过基于Shapley值的公平特征归因方法,从整体视角处理数据掩码问题,不同于现有文献大多关注数据集层面的风险-效用权衡,本框架在特征层面解决该权衡问题。此外,该框架对数据掩码方法、统计与机器学习方法、数据效用及披露风险评估指标均无特定要求。实验结果表明,所提方法能在保留数据效用的同时有效降低披露风险。
英文摘要
Despite its many benefits, widespread access to individuals' personal data also causes severe privacy concerns for consumers, companies, and policymakers. This study proposes a novel framework that adapts the Shapley-value-based feature attribution approach to the problem domain of data privacy by capturing the two crucial dimensions of data privacy---disclosure risk and data utility. Our proposed framework takes a holistic view of data masking through a fair feature attribution approach based on Shapley values. Different from the existing literature that mostly focuses on the risk-utility tradeoff at the dataset level, the proposed framework addresses the tradeoff at the feature level. Furthermore, the proposed framework is agnostic to data masking methods, statistical and machine learning methods, and data utility and disclosure risk evaluation metrics. Experimental results show that our proposed method can effectively reduce disclosure risk while preserving data utility.