arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向联邦检索增强生成的私有任意时刻选择性风险认证:保证与经验极限

Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation: Guarantees and Empirical Limits

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

arXiv 2608.07913首次发表:更新:

发表机构

Iowa State University(爱荷华州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Fed-SRC这一适用于联邦RAG的私有任意时刻选择性风险认证方法,经实验验证其在多数场景下可满足风险保证,性能优于对比方法。

AI 中文摘要

选择性风险证书保证被接受的输出满足声明的错误目标。我们开发了Fed-SRC,这是一种与分数无关的证书,适用于联邦、差分隐私、自适应监控的检索增强生成(RAG)。客户端仅发布经过高斯扰动的分数和损失直方图。以记录索引和噪声方差索引的鞅共同约束所有注册阈值和轮次上的目标风险对比与被接受的质量,允许可预测的招募、退出、阈值选择和可选停止。范围为1的总变差项将校准混合物转移到声明的部署混合物。本研究的贡献在于这种私有、联邦化、任意时刻的组合,而非单独的对比统计量或接受下限。实验中,在所有评估单元、隐私级别或策略中均未出现同时边界违反情况。操作能力取决于分数和总体:主要目标r*=0.10从未通过认证,在RAGTruth上的次要目标r*=0.20也从未通过认证,而在HaluEval问答任务上,其在200次非私有试验中全部通过认证,保留风险低于目标。未经私有化的非私有证书在200次试验中有146至198次违反其边界。作为探索性比较,我们还评估了一种私有赌资启发式方法,我们未确立其e过程有效性,该启发式方法在ε≤4时停止认证,而Fed-SRC仍可认证,不过认证消耗的流事件约为唯一校准项的30倍。

英文摘要

Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate for federated, differentially private, adaptively monitored retrieval-augmented generation. Clients release only Gaussian-perturbed score and loss histograms. Record-indexed and noise-variance-indexed martingales jointly bound target-risk contrast and accepted mass over all registered thresholds and rounds, permitting predictable recruitment, dropout, threshold selection, and optional stopping. A range-one total-variation term transfers the calibration mixture to a declared deployment mixture. The contribution is this private, federated, anytime combination, rather than the contrast statistic or acceptance floor individually. Empirically, no simultaneous-bound violation occurs in any evaluated cell, privacy level, or policy. Operational power depends on the score and population: the primary target r*=0.10 never certifies, and on RAGTruth the secondary target r*=0.20 never certifies either, whereas on HaluEval question answering it certifies in all 200 non-private trials, with held-out risk below the target. Naively privatized non-private certificates violate their bounds in 146 to 198 of 200 trials. As an exploratory comparison, we also evaluate a private betting-capital heuristic for which we do not establish e-process validity. This heuristic stops certifying at epsilon <= 4, where Fed-SRC still certifies. Certification nevertheless consumes roughly 30 times more stream events than unique calibration items.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑