可学习的随机化作为对抗自适应优化器的承诺
Learnable Randomization as Commitment Against Adaptive Optimizers
浏览论文内容
中文总结 AI 辅助
针对自适应优化器,提出可学习的随机化混合策略(UNOP)作为承诺,在定价和推荐中保留用户剩余,优于贪婪策略,但在特定条件下应关闭。
中文摘要 AI 辅助
定价页面可以将标价逐步提高到买家仍然接受的最后金额,推荐系统可以扣留更好的商品而推荐一个勉强可接受的推广商品,分类器可以在申请人改变其特征后移动其决策边界。系统预测响应,然后选择服务于自身目标的菜单,因此用户截止值以上的剩余被攫取。采用单一最佳行动会公开该截止值,而对用户永远不会采取的行动施加噪声则会浪费收益,并让平台认为更差的菜单仍然可以接受。我们研究不可预测的近似最优策略(UNOP),该策略在保持个体理性的近似最优行动上均匀混合。这种混合是对响应的一种承诺。在有限价格网格上,当最佳确定性需求价格严格优于随机化区间时,已知曲线的卖家会将价格设定在区间之下,而发生的购买是确定性的。知道该曲线与预测下一次抽取并不相同。这种混合可以被学习,优化器可以匹配其最佳响应,而用户的收益保持更高,因为混合改变了被针对的行动。在定价和策略感知推荐中,当平台针对曲线进行优化且不止一个行动可接受时,这比贪婪策略留下更多剩余。在质量排序、单元素近似最优集、错误的效用估计或短视探索者的情况下,这种收益消失。这也是当另一方试图合作时应关闭混合的地方。
英文摘要
A pricing page can walk the posted price up to the last amount a buyer still accepts, a recommender can hold back a better item for a barely acceptable promoted one, and a classifier can shift its boundary once applicants change their features. The system predicts the response and then picks the menu that serves its own objective, so the surplus above the user's cutoff is taken. Playing the single best action publishes that cutoff, while noise on actions the user would never take throws away payoff and teaches the platform that a worse menu is still acceptable. We study unpredictable near-optimal policies (UNOP), which mix uniformly on near-best actions that remain individually rational. The mixture is a commitment about the response. On a finite price grid, when the best sure-demand price strictly out-earns the randomized band, a seller who already knows the curve posts below the band, and the purchase that occurs is deterministic. Knowing that curve is not the same as predicting the next draw. The mixture can be learned and the optimizer can match its best response, while the user's payoff stays higher because the mixture changes which action is targeted. In pricing and in policy-aware recommendation this leaves more surplus than greedy play when the platform optimizes against the curve and more than one action is acceptable. The gain goes away under quality ranking, a singleton near-optimal set, a wrong utility estimate, or a short-horizon explorer. That is also where mixing should be turned off if the other side is trying to cooperate.
发表机构
- The University of Hong Kong(香港大学)
- The University of Sydney(悉尼大学)
- University of Electronic Science and Technology of China(电子科技大学)
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。