ARENA:大型音频语言模型的自动化红队测试
ARENA: Automated Red-Teaming for Large Audio Language Models
浏览论文内容
中文总结 AI 辅助
本文提出ARENA闭环框架,基于2000个文本-音频数据集训练控制器,在520个AdvBench目标上对多个LALMs实现高FDR/PSR,验证了其在音频红队测试中的有效性。
中文摘要 AI 辅助
大型音频语言模型(LALMs)支持通过语音、音乐和环境声音与语言模型交互,但也引入了仅用文本红队测试难以暴露的安全风险。本文研究基于音频的自动化红队测试,其中文本查询单独存在时保持安全,而文本-音频联合输入会诱导有害目标行为。我们提出ARENA,这是一个闭环框架,使用独立的2000个案例的文本-音频数据集训练控制器;MD-Judge提供训练奖励和自适应搜索反馈,而单独的非自适应Llama Guard 3评估器仅对最终结果进行标注。在520个保留的AdvBench目标上,ARENA在Audio Flamingo 3、Qwen2-Audio、MiMo-Audio和GPTAudio上分别实现了87.9/100.0%、71.5/96.3%、68.1/100.0%和75.4/98.5%的FDR/PSR; ablation实验表明,基于反馈的优化和音频变体搜索大幅提升了攻击发现能力。
英文摘要
Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that is difficult to expose with text-only red-teaming. We study automated audio-grounded red-teaming, where a text query must remain safe in isolation while the joint text-audio input induces harmful target behavior. We propose ARENA, a closed-loop framework that trains a controller on an independent 2,000case text-audio dataset. MD-Judge supplies training rewards and adaptive search feedback, while a separate, non-adaptive Llama Guard 3 evaluator alone labels final outcomes. On 520 held-out AdvBench objectives, ARENA achieves FDR/PSR of 87.9/100.0%, 71.5/96.3%, 68.1/100.0%, and 75.4/98.5% on Audio Flamingo 3, Qwen2-Audio, MiMo-Audio, and GPTAudio, respectively. Ablations show that feedback-based refinement and audio-variant search substantially improve attack discovery.