arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PrivCert:在差分隐私下认证语句支持度

PrivCert: Certifying Statement Support under Differential Privacy

Tsubasa Takahashi, Takumi Hiraoka

arXiv 2609.38934首次发表:更新:

发表机构

Acompany Co., Ltd.(Acompany有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PrivCert提出隐私保护报告框架,通过证书和弃权决策显式认证语句支持度,弥补差分隐私文本生成的证据缺口,实验表明其优于基线。

AI 中文摘要

差分隐私(DP)文本生成可以保护个体记录,但仅凭隐私并不能说明发布的语句携带了关于底层数据的何种证据。我们将此识别为证据缺口:一份私有报告可能包含看似合理的声明,却未指明这些声明是否得到私有数据集的强有力支持。我们提出了PrivCert,一个用于隐私保护报告的框架,通过隐私保护证书和“发出或弃权(不执行)”决策使语句支持度明确化。作为一个典型实例,PrivCert-PF(提议与过滤)将数据无关的候选发现与私有支持度认证分离,仅发出那些支持度通过私有证据测试的语句。我们为该框架提供了理论基础,通过刻画差分隐私下隐式证据的局限性,推导出单语句认证的尖锐隐私-诚实边界,并为细粒度多语句认证建立了最坏情况成本。在合成任务以及TAB、WildChat和Yelp上的实验表明,显式认证保持了较低的无支持发出率,而自由文本差分隐私基线在相同的声明支持语义下经常产生低支持度的声明。我们进一步展示了PrivCert契约可以通过直方图、稀疏向量和高斯机制实现,并使用差分隐私合成数据说明一个重要边界:私有代理中的支持度并不自动认证原始数据中的支持度。综合来看,这些结果将隐私保护报告定位为一个证据设计问题:不仅是如何生成私有文本,还有私有报告能对其底层数据证实什么。

英文摘要

Differentially private (DP) text generation can protect individual records, but privacy alone does not specify what evidence a released statement carries about the underlying data. We identify this as an evidence gap: a private report may contain plausible claims without indicating whether they are strongly supported by the private dataset. We introduce PrivCert, a framework for privacy-preserving reporting that makes statement support explicit through privacy-preserving certificates and emit-or-abstain decisions. As a canonical instantiation, PrivCert-PF (Proposal-and-Filter) separates data-independent candidate discovery from private support certification, emitting only statements whose support passes a private evidence test. We provide theoretical grounding for this framework by characterizing the limits of implicit evidence under DP, deriving a sharp privacy--honesty frontier for single-statement certification, and establishing a worst-case cost for fine-grained multi-statement certification. Experiments on synthetic tasks and TAB, WildChat, and Yelp show that explicit certification maintains low unsupported emission, while free-text DP baselines frequently produce low-support claims under the same declared support semantics. We further show that the PrivCert contract can be realized with histogram, sparse-vector, and Gaussian mechanisms, and use DP synthetic data to illustrate an important boundary: support in a private proxy does not automatically certify support in the original data. Together, these results position privacy-preserving reporting as an evidence-design problem: not only how to generate private text, but what a private report can substantiate about its underlying data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑