arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.31105cs.AIcs.CL

BLOOM-WILT:用于自动化大语言模型审计中行为 elicitation 的 Logit Tilting

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing

Adrians Skapars, Edoardo Manino

首次发表
浏览论文内容

中文总结 AI 辅助

BLOOM-WILT是一种自动化大语言模型审计流程,通过自适应重新加权解码提升采样效率,在4个模型8种行为的评估中多数优于基线,还推翻了此前的模型安全排名。

中文摘要 AI 辅助

部署后的语言模型用户常会遇到测试几乎无法发现的行为,因为部署阶段的交互量远超任何评估可模拟的规模。自动化审计工具让测试成本降低且可灵活覆盖几乎所有指定行为,但这类工具缺乏优化压力,采样效率低下。为解决该缺陷,我们提出BLOOM-WILT,这是一个完整的审计流程,无需训练成本,也无需访问目标模型的下一个 token 分布之外的内容,即可引出罕见行为的自然多轮实例。在输入侧,WILT 的审计模型会在多轮对话中调整策略,从之前已评分的交互中学习;在输出侧,WILT 会基于引出提示词,利用目标模型自身的分布自适应地重新加权解码过程,使得与目标行为相关的生成内容,在未加提示时与其他内容概率相等,加提示后会被优先采样。我们在4个目标模型和8种行为上对WILT进行评估,结果显示,在32种设置中,WILT有30种表现优于基线审计工具,还推翻了此前的模型安全排名。当从Qwen3.5-4B引出自残鼓励内容时,WILT将平均行为出现率从51%提升至100%,在计算资源匹配的情况下,优于我们移植到同一流程中的所有引出方法,且不会将输出概率压低至基线以下。

英文摘要

Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deployment puts the model through orders of magnitude more interactions than any evaluation can simulate. Automated auditors make testing cheap to scale and flexible enough to cover almost any specified behaviour, yet their lack of optimisation pressure makes them sample-inefficient. To address this shortcoming, we introduce BLOOM-WILT, a full auditing pipeline that elicits natural multi-turn instances of rare behaviours, without training cost or access beyond the target's next-token distribution. On the input side, WILT's auditor model revises its conversational strategy across rounds, learning from previous scored interactions. On the output side, WILT adaptively reweights the target's decoding using the model's own distribution conditioned on an elicitation prompt, so that behaviour-relevant generations are sampled ahead of others it finds equally probable when unprompted. We evaluate WILT across 4 target models and 8 behaviours, where it beats the baseline auditor in 30 of the 32 settings and overturns the previous model safety rankings. WILT raises average behaviour presence from 51% to 100% when eliciting self-harm encouragement from Qwen3.5-4B, beating every elicitation method we port into the same pipeline at matched compute, without pushing output probability below the baseline's.

发表机构

  • University of Manchester(曼彻斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑