arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16283cs.LG

用于聚合洞察生成的差分隐私语义计划

Differentially Private Semantic Plans for Aggregate Insight Generation

发表机构哈佛大学
查看机构详情
  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

Behrooz Razeghi

首次发表
浏览论文内容

中文总结 AI 辅助

DP-SPIN是一个可信策展人框架,通过将记录映射为语义草图并发布差分隐私语义计划,实现独立于受保护数据的语义概念的聚合测量与摘要生成。

中文摘要 AI 辅助

URANIA为数据相关聚类的摘要提供端到端差分隐私(DP)。然而,其聚类-关键词发布并不直接为独立于受保护语料库定义的语义概念提供集合范围的聚合。记录可能表达多个概念,表达相同概念的记录可能被分配到不同的聚类,并且聚类标识在不同分析中不必对应。因此,聚类级统计量不能直接提供跨集合或重复分析中预定义概念的可比较测量。我们引入了DP-SPIN,一个可信策展人框架,用于对独立于受保护目标记录固定的语义概念进行聚合测量和摘要。每条记录被映射到这些概念上的一个有界稀疏非负向量,其总和形成语义草图。一个差分隐私机制发布一个包含被接纳概念和噪声质量的语义计划;通过后处理获得归一化的语义支持值和支持箱。对于用户级隐私,每个用户的聚合贡献被裁剪到固定界限。语言模型仅接收计划和固定的解码指令,而公共验证器检查概念提及、报告值、比较和排名声明是否与发布的计划一致。最终摘要通过后处理实现差分隐私。我们在添加/删除和替换邻接关系下建立了记录级和用户级DP保证。我们在CFPB投诉叙述、Amazon All Beauty评论和Yelp餐厅评论上以记录级隐私评估DP-SPIN,并在Amazon和Yelp上以用户级隐私评估。我们将DP-SPIN与非私有计划和摘要参考、DP关键词和类别直方图基线以及具有固定公共关键词词汇表的URANIA风格基线进行比较。

英文摘要

\texttt{URANIA} provides end-to-end differential privacy (DP) for summaries of data-dependent clusters. However, its cluster--keyword release does not directly provide collection-wide aggregates for semantic concepts defined independently of the protected corpus. Records may express several concepts, records expressing the same concept may be assigned to different clusters, and cluster identities need not correspond across analyses. Consequently, cluster-level statistics do not directly provide comparable measurements of predefined concepts across collections or repeated analyses. We introduce \texttt{DP-SPIN}, a trusted-curator framework for aggregate measurement and summarization over semantic concepts fixed independently of the protected target records. Each record is mapped to a bounded sparse nonnegative vector over these concepts, whose sum forms a semantic sketch. A differentially private mechanism releases a semantic plan containing admitted concepts and noisy masses; normalized semantic-support values and support bins are obtained by post-processing. For user-level privacy, each user's aggregate contribution is clipped to a fixed bound. The language model receives only the plan and fixed decoding instructions, while a public verifier checks concept mentions, reported values, comparisons, and rank claims against the released plan. The final summary is differentially private by post-processing. We establish record- and user-level DP guarantees under add/drop and replacement adjacency. We evaluate \texttt{DP-SPIN} under record-level privacy on CFPB complaint narratives, Amazon All Beauty reviews, and Yelp restaurant reviews, and under user-level privacy on Amazon and Yelp. We compare \texttt{DP-SPIN} with non-private plan and summary references, DP keyword and category histogram baselines, and a \texttt{URANIA}-style baseline with a fixed public keyword vocabulary.

↑