AEGIS:基于内部信号的音频内生防护,抵御大型音频语言模型越狱攻击
AEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks
- National Taiwan University(国立台湾大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大型音频语言模型的越狱攻击,提出AEGIS防御方法,通过中层风险门控激活安全适配器,将平均不安全率从17.9%降至0.4%,实现鲁棒拒绝。
AI中文摘要:
大型音频语言模型(LALMs)扩展了语言模型处理和理解音频的能力,但也使其面临异构音频越狱攻击的风险。我们探究了成功的越狱是反映了未能识别有害意图,还是在识别之后发生的失败。逐层探测揭示了后者:风险相关信息仍可从中间表示中解码,但内部风险信号未能在后续层处理中转化为拒绝。我们将这一差异识别为风险-拒绝差距。基于此发现,我们提出了AEGIS,一种检测后干预的防御方法,其中层风险门控选择性地激活下游安全适配器。在六个LALMs和三个异构音频越狱基准测试中,AEGIS将平均不安全率从17.9%降至0.4%,同时仅导致良性输入上过度拒绝的轻微增加。这些结果确立了选择性内部干预作为增强LALMs鲁棒拒绝的有效途径。代码可在该https URL获取。
英文摘要:
Large audio-language models (LALMs) expand language models to process and interpret audio, but also expose them to heterogeneous audio jailbreaks. We ask whether successful jailbreaks reflect failures to recognize harmful intent or failures occurring after such recognition. Layer-wise probing reveals the latter: risk-related information remains decodable from intermediate representations, yet the internal risk signal fails to translate into refusal in later-layer processing. We identify this discrepancy as the risk-to-refusal gap. Building on this finding, we propose AEGIS, a detect-then-intervene defense whose mid-layer risk gate selectively activates downstream safety adapters. Across six LALMs and three heterogeneous audio jailbreak benchmarks, AEGIS reduces the average unsafe rate from 17.9% to 0.4%, while causing only a marginal increase in over-refusal on benign inputs. These results establish selective internal intervention as an effective path toward more robust refusal in LALMs. The code is available at https://github.com/azzzzliao/aegis-audio-defense.