arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31142cs.CRcs.AIcs.CL

JevAdvBench:用于校准决策强化学习模型的基准与黑盒攻击

JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models

Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang, Leo Yu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对RLCD模型提出首个对抗基准JevAdvBench,通过对比模型自身干净决策评估攻击,发现附加意见可翻转12.1%决策,建议将状态视为不可信输入。

中文摘要 AI 辅助

使用强化学习进行校准决策(RLCD)训练的模型(如Jev)会针对输入(即状态)回答一个类型化问题,输出概率、选择或分数,软件无需人工阅读即可根据答案采取行动。其鲁棒性尚未得到衡量:对抗性基准评估的是模型生成或执行的内容,而类型化模型不生成任何内容,即使被操纵也会返回格式良好的答案。测量也很困难,因为相同请求可能返回不同答案,大多数可用标签来自模型本身,且API在视线之外预处理每个请求。我们的关键思想是将每个受攻击的决策与模型自身的干净决策进行比较,而不是与标签比较,并通过相同重运行引起的变化来解读该决策。基于此,我们引入了JevAdvBench,据我们所知,这是首个针对RLCD模型的对抗性基准,包含66个场景下的812个类型化问题,以及一个包含9,744个单编辑变体的黑盒攻击套件,每个变体编辑请求的一个部分,并通过计费输入令牌确认编辑已到达模型。在jev-1.13.0上,改写保持在重运行基线的1.2个百分点以内,模式之外的字段从未到达模型。相比之下,附加到状态的一条未经核实的意见翻转了12.1%的决策,与最强的注入命令(10.1%)在统计上相当,并将38%的自信答案推至0.8置信度阈值以下,该阈值将其路由至人工审查。因此,基于RLCD模型构建的应用程序应将状态视为不可信的、有争议的输入。项目网站:此https URL

英文摘要

Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can return different answers, most available labels come from the model itself, and the API preprocesses each request out of view. Our key idea is to score each attacked decision against the model's own clean decision rather than against labels, and to read it against the change caused by an identical re-run. Building on this, we introduce JevAdvBench, to our knowledge the first adversarial benchmark for RLCD models, with 812 typed questions over 66 scenarios, and a black-box attack suite of 9,744 single-edit variants that each edit one part of a request, with billed input tokens confirming that the edit reached the model. On jev-1.13.0, rewording stays within 1.2 percentage points of the re-run baseline, and fields outside the schema never reach the model. In contrast, one unverified opinion appended to the state flips 12.1% of decisions, statistically tied with the strongest injected command (10.1%), and pushes 38% of confident answers below the 0.8 confidence threshold that routes them to human review. Applications built on RLCD models should therefore treat the state as untrusted, argued input. Project website: https://JevAdvBench.github.io/JevAdvBench/

发表机构

  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
  • School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
  • Huazhong University of Science and Technology(华中科技大学)
  • Griffith University(格里菲斯大学)
  • Changsha University of Science and Technology(长沙理工大学)
  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑