AI 中文总结
研究差分隐私下的最大和多样化问题,提出基数和拟阵约束下的差分隐私算法,实现近最优效用保证,设计更高效算法,实验表明该方法在强隐私保证下效用与非私密基线相当且显著提高执行时间。
AI 中文摘要
结果多样化对于生成信息丰富、无冗余的数据摘要和查询输出至关重要。尽管其各种形式在一系列数据驱动学科中得到了广泛研究,但现有方法未能解决基础数据敏感时出现的隐私问题。本文研究差分隐私下的结果多样化,聚焦最大和多样化(MSD)问题,提出基数和拟阵约束下的差分隐私算法,实现近最优效用保证,设计更高效算法,在真实数据集上实验表明该方法在强隐私保证下效用与非私密基线相当且显著提高执行时间。
英文摘要
Result diversification is crucial for generating informative, non-redundant data summaries and query outputs. Although its various formulations have been extensively studied across an array of data-driven disciplines, existing methods fail to address the privacy concerns that arise when the underlying data is sensitive. In this work, we initiate the study of result diversification under differential privacy, focusing on the max-sum diversification (MSD) problem, a widely adopted model with the objective of maximizing a linear combination of a submodular function, quantifying relevance, and the sum of pairwise distances between selected items, quantifying diversity. We propose differentially private algorithms for MSD under both cardinality and matroid constraints, achieving nearly optimal utility guarantees. At the same time, we design more efficient algorithms that maintain strong guarantees. Notably, the proposed algorithms are faster than existing non-private methods, making them appealing even in non-private settings. Experimental evaluations on real-world datasets demonstrate that the proposed approach achieves utility comparable to that of non-private baselines even under strong privacy guarantees, and significantly improves execution times for cardinality constraints.
CommentsVLDB 2026