Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
将偏见转化为漏洞:基于Bandit引导的LLM裁判风格操纵攻击
机构 * School of Computing, National University of Singapore, Singapore(新加坡国立大学计算机学院) ; Nanyang Technological University, Singapore(南洋理工大学)
AI总结 提出BITE黑盒对抗框架,将风格编辑选择建模为上下文Bandit问题,通过LinUCB策略自适应选择编辑以误导LLM裁判并人为提高评分,攻击成功率超65%。
Comments Accepted to the Forty-Third International Conference on Machine Learning (ICML 2026)