arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越奖励抑制:有界奖励下热启动多臂老虎机的近最优离线攻击

Beyond Reward Suppression: Near-Optimal Offline Attacks on Warm-Start Bandits with Bounded Rewards

Qirun Zeng, Manhin Poon, Xiangxiang Dai, Qixin Zhang, Jinhang Zuo

arXiv 2610.10000首次发表:更新:

发表机构

City University of Hong Kong; The Chinese University of Hong Kong; Nanyang Technological University(香港城市大学; 香港中文大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究热启动多臂老虎机的离线对抗攻击,证明目标提升的必要性,并提出实现最优次线性成本的攻击方法,扩展至多种算法并经实验验证。

AI 中文摘要

针对多臂老虎机的对抗性攻击旨在以较小的攻击成本误导学习器选择目标臂。现有攻击通常通过抑制非目标臂来实现这一目标。然而,在实践中,诸如虚假评论之类的操纵往往直接提升目标物品。我们通过有界奖励下热启动多臂老虎机的离线攻击来研究这一差距,其中攻击者只能在部署前向热启动历史中注入有效的动作-奖励对。我们表明,目标提升并非仅仅是启发式方法:当目标臂位于较低奖励边界附近时,任何针对UCB的、使其在几乎所有在线轮次中被选中的最优成本攻击,都必须将其成本的非可忽略部分分配给目标臂。然后,我们设计了一种实现最优次线性成本的攻击,并刻画了其在目标提升与非目标抑制之间的成本分配。我们进一步将攻击扩展到汤普森采样、ε-贪心以及更广泛的一类多臂老虎机算法。在真实世界和合成数据上的实验验证了我们攻击的有效性。

英文摘要

Adversarial attacks on bandits aim to mislead a learner toward a target arm while keeping the attack cost small. Existing attacks typically achieve this by suppressing non-target arms. In practice, however, manipulation such as fake reviews often directly promotes the target item. We study this gap through bounded offline attacks on warm-start bandits, where an attacker can inject only valid action-reward pairs into the warm-start history before deployment. We show that target promotion is not merely a heuristic: when the target arm lies near the lower reward boundary, any order-optimal-cost attack against UCB that makes it selected in nearly all online rounds must allocate a nonvanishing fraction of its cost to the target arm. We then design an attack that achieves the optimal sublinear cost and characterize its allocation between target promotion and non-target suppression. We further extend the attack to Thompson Sampling, $ε$-greedy, and a broader class of bandit algorithms. Experiments on real-world and synthetic data validate the effectiveness of our attacks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑