只需询问 Jev:用于校准决策的强化学习作为 AI 对齐失败的零样本检测器
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
浏览论文内容
中文总结 AI 辅助
提出 Jev 模型,通过强化学习训练实现校准决策,在单次调用中检测多种 AI 对齐失败,零样本 AUROC 达 0.886,成本比 LLM 评判器低 63 倍。
中文摘要 AI 辅助
对齐失败检测器用于筛选部署的语言模型并评分对齐基准。大多数检测器是生成式评判器,对每个标准进行一次解码,而像 Llama Guard 这样读取令牌概率的分类器,每次调用仅对一个固定标签进行评分。Jev 是一个通过强化学习训练用于校准决策(RLCD)的模型,能在单次调用中针对一个输入回答多个类型化问题,并提供校准概率。其是否能检测对齐失败尚未被评估。我们提出了 RLCDAlignBench,该基准在十种对齐失败上对 Jev 进行评测:谄媚、越狱、欺骗、提示注入、幻觉、隐私侵犯、社会偏见、奖励黑客、隐瞒不确定性以及追求权力。它涵盖 44 个基准和五个目标模型,由每个基准的评分器进行标注,并在其中两个基准上由人类标注。许多此类失败是关系型的,即相对于参考(如用户的信念或注入的指令)进行定义,而仅凭响应本身无法揭示这些参考。因此,我们的关键思路是分别变化对 Jev 的提问内容与其所见内容:一方面改变问题的措辞和答案类型,另一方面改变输入的字段。一个单一的通用问题在零样本情况下达到 0.886 的中位 AUROC,并在大多数基准上超越监督基线。问题措辞影响甚微,而上下文影响更大,主要通过编码标签的字段起作用。Jev 与参考评分器在人类标签上的一致性相当,揭示了现有基准中的标签缺陷,并且其成本比 LLM 评判器评分器低 63 倍。代码和数据:此 https URL。
英文摘要
Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not been measured. We present RLCDAlignBench, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. It spans 44 benchmarks and five target models, labelled by each benchmark's scorer and, on two, by humans. Many of these failures are relational, defined against a reference, such as the user's belief or an injected instruction, that the response alone does not reveal. Our key idea is therefore to vary what Jev is asked separately from what it sees: the question's wording and answer type on one side, the fields of the input on the other. A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks. Question wording matters little, while context matters more, mostly through fields that encode the label. Jev matches the reference scorer's agreement with human labels, surfaces label defects in existing benchmarks, and costs 63x less than LLM-judge scorers. Code and data: https://github.com/sumleo/RLCDAlignBench.
发表机构
- Griffith University(格里菲斯大学)
- Nanyang Technological University(南洋理工大学)
- UNSW(新南威尔士大学)
- Deakin University(迪肯大学)
- George Mason University(乔治梅森大学)
- Wake Forest University(维克森林大学)
机构由 AI 辅助整理,请以论文原文为准。