AI 中文总结
该研究针对在线实验中第三方推断消费者群体的问题,提出含两种策略的隐私预算框架,经78个实验验证,动态策略适配复杂实验,优化预算分配可提升学习性能。
AI 中文摘要
企业使用多臂老虎机(multi-armed bandits)开展在线实验,在平衡探索与利用的同时为消费者推送个性化内容。然而,第三方可通过观察消费者收到的横幅、广告或推荐内容,推断其潜在群体。为控制此类推断,我们提出一种隐私风险预算,企业可预先设定该预算,利用差分隐私(differential privacy)限制第三方的信念更新。为使用该隐私风险预算,我们提出两种策略:常数隐私风险策略,以及根据不同访问者差异化使用隐私风险的动态隐私风险策略。我们在网站设计和推荐系统两个应用场景中,研究这两种策略下隐私风险预算对实验性能的影响。针对两种策略,我们通过分析得出可最优平衡探索与利用的隐私风险预算。随后,我们将实验级隐私风险预算的概念扩展至企业级隐私风险预算,并在包含78个实验的实证场景中应用该预算。研究发现,动态策略在更长、更复杂的实验中尤为有价值,且优化企业级隐私风险预算在各实验间的分配可显著提升学习性能。
英文摘要
Firms perform online experiments with multi-armed bandits to personalize what consumers are shown while balancing exploration and exploitation. However, third-parties can infer consumers' underlying segments from observing which banners, ads, or recommendations consumers receive. To control this inference, we propose a privacy risk budget that firms can set ex ante to bound such third party belief updating using differential privacy. To spend this privacy risk budget, we propose two strategies: a constant privacy risk strategy and a dynamic privacy risk strategy that spend privacy risk differently across visitor. We study how privacy risk budgets affect experimentation performance in two applications--website design and a recommendation system--under these strategies. For both strategies, we analytically find privacy risk budgets that optimally balance exploration and exploitation. We then extend the idea of an experiment-level privacy risk budget to a firm-wide privacy risk budget. We apply this firm-wide privacy risk budget in an empirical setting with 78 experiments. We find that the dynamic strategy is particularly valuable in longer and more complex experiments, and that optimizing the allocation of a firm-wide privacy risk budget across experiments substantially improves learning performance.