发表机构
Institute of Data Science; National University of Singapore; Department of Electrical and Computer Engineering; Department of Mathematics(数据科学研究所; 新加坡国立大学; 电气与计算机工程系; 数学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出AO-Pri-BAI算法,在纯ε-差分隐私下实现固定预算最佳臂识别的渐近最优指数衰减率,并通过拉普拉斯树机制和最小-最大采样设计达到隐私感知基准,数值实验验证其优势。
AI 中文摘要
差分隐私下的最佳臂识别是一个纯探索问题,其中必须同时实现统计效率和隐私保护。我们研究了在纯ε-差分隐私下老虎机的固定预算最佳臂识别,其中学习者在规定的采样预算后必须推荐一个臂,同时保护完整的记录。我们证明了错误概率的最优指数衰减率上界由一个依赖于实例的隐私感知传输指数给出,该指数不同于Jourdan和Azize [2025]在固定置信分析中用于刻画停止时间的类似量。在此指数的指导下,我们提出了AO-Pri-BAI,一种自适应算法,通过拉普拉斯树机制维护私有的运行估计,并通过硬替代方案与臂分配之间的最小-最大交互来学习采样设计。我们证明了AO-Pri-BAI满足纯ε-差分隐私。我们还建立了AO-Pri-BAI的失败概率指数与隐私感知基准相匹配。数值研究表明,即使在非渐近设置中,AO-Pri-BAI在各种实例上也优于基准算法,补充了理论分析。
英文摘要
Best arm identification under differential privacy is a pure-exploration problem in which both statistical efficiency and privacy protection must be achieved simultaneously. We study fixed-budget best arm identification for bandits under pure $ε$-differential privacy, where the learner must recommend an arm after a prescribed sampling budget while protecting the full transcript. We prove that the optimal exponential decay rate of the error probability is upper bounded by an instance-dependent privacy-aware transportation exponent that differs from the analogous quantity used to characterize the stopping time in fixed-confidence analysis by Jourdan and Azize [2025]. Guided by this exponent, we propose AO-Pri-BAI, an adaptive algorithm that maintains private running estimates through Laplace-tree mechanisms and learns a sampling design through a min--max interaction between hard alternatives and arm allocations. We prove that AO-Pri-BAI satisfies pure $ε$-differential privacy. We also establish that the exponent of the failure probability of AO-Pri-BAI matches the privacy-aware benchmark. Numerical studies show that even in the non-asymptotic setting, AO-Pri-BAI outperforms benchmark algorithms on various instances, complementing the theoretical analyses.
CommentsAccepted to NeurIPS 2026