arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25041cs.IRcs.CRcs.CVcs.LG

ScoreShield:相似度分数的差分隐私发布

ScoreShield: Differentially Private Release of Similarity Scores

Behrooz Razeghi, Parsa Rahimi

首次发表
浏览论文内容

中文总结 AI 辅助

研究生物识别等应用中相似度分数隐私保护问题,提出ScoreShield先扰动后投影机制,添加校准高斯噪声并投影到可行性集,满足(ε,δ)-DP,给出效用保证,在多任务中评估,改进了风险的n依赖性。

中文摘要 AI 辅助

越来越多的应用,如生物识别和检索增强生成(RAG),依赖于文本、图像或音频向量嵌入之间计算的余弦相似度分数。这些系统通过API返回相似度分数用于排名和验证,但可能泄露个体记录信息并引发成员推理攻击。虽然差分隐私(DP)提供了量化攻击风险的原则性度量,但简单应用DP机制(如给向量条目添加独立同分布高斯噪声)在给定隐私约束下会导致过度失真(即低效用),且与发布分数数量的扩展性差。我们提出ScoreShield,一种先扰动后投影的机制,添加根据所选分数发布机制的全局敏感度校准的高斯噪声,然后将结果投影到有效余弦对象的可行性集上。ScoreShield在发布相似度分数向量和Gram矩阵时满足(ε,δ)-DP。我们为风险分析中使用的精确Frobenius度量投影提供效用保证,并证明用于大规模Gram发布的实际平均交替投影求解器收敛到可行性。对于记录级替换邻接下的全成对余弦Gram发布,精确投影界将平方Frobenius风险的n依赖性从朴素高斯基线的Θ(n^3)改善到固定隐私参数下的O(n^2),在低秩Gram时有更尖锐的局部界。我们在RAG、人脸识别、语义检索、图像相似度和推荐系统任务中评估了该机制。

英文摘要

A growing number of applications, such as biometrics and retrieval-augmented generation (RAG), rely on cosine similarity scores computed between vector embeddings of text, images, or audio. These systems return similarity scores through their APIs for ranking and verification. However, such releases can leak information about individual records and enable membership inference attacks. While differential privacy (DP) provides a principled metric for quantifying attack risks, naïve application of DP mechanisms---such as adding i.i.d. Gaussian noise to vector entries---leads to excessive distortion (i.e., low utility) at a given privacy constraint that scales poorly with the number of released scores. We propose \textsc{ScoreShield}, a perturb-then-project mechanism that adds Gaussian noise calibrated to global sensitivity of the chosen score release regime and then projects the result onto the feasibility set of valid cosine objects. \textsc{ScoreShield} satisfies $(\varepsilon,δ)$-DP for releasing similarity score vectors and Gram matrices. We provide utility guarantees for the exact Frobenius metric projection used in the risk analysis, and prove convergence to feasibility for the practical averaged alternating-projection solver used for large-scale Gram releases. For full pairwise cosine Gram release under record-level replacement adjacency, the exact-projection bound improves the $n$-dependence of squared Frobenius risk from $Θ(n^3)$ for the naïve Gaussian baseline to $\mathcal{O}(n^2)$ for fixed privacy parameters, with sharper local bounds at low-rank Grams. We evaluate the mechanism across RAG, face recognition, semantic retrieval, image similarity, and recommender-system tasks.

发表机构

  • School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院)
  • School of Engineering, École Polytechnique Fédérale de Lausanne (EPFL)(洛桑联邦理工学院工程学院)

机构由 AI 辅助整理,请以论文原文为准。

↑