发表机构
Beijing University of Posts and Telecommunications; Institute of Information Engineering, Chinese Academy of Sciences; Nanyang Technological University; JIUTIAN Research; Tencent ARC Lab; Chongqing University of Posts and Telecommunications(北京邮电大学; 中国科学院信息工程研究所; 南洋理工大学; 中移九天; 腾讯ARC实验室; 重庆邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出不可听红队测试方法ILL评估LALMs的低频安全风险,发现其可降低LALMs准确率达67个百分点,同时提出DRG防护方法可提升受攻击后的准确率,明确了LALMs此前被忽视的安全风险。
AI 中文摘要
大型音频语言模型(LALMs)在理解各类音频输入方面展现出强大能力,其中包括人类不可听的低频信号,这类信号仍能进入模型并影响其生成。然而,这类低频输入对LALMs的实际影响在很大程度上尚未被探索。本文提出间歇性低频锁定(Intermittent Low-Frequency Lockout,ILL),一种在黑盒环境下使用通用波形模板评估该风险的不可听红队测试方法。ILL采用句子注意力尺度估计确定活跃区间,并利用频率混淆迁移从语料库频谱变化构建具有连续相位的低频波形。为缓解该风险,本文提出分布重查询防护(Distributional Requery Guard,DRG)以检测低频分布偏移,并有条件地请求第二次录音用于语义恢复。在6个LALMs和多项音频理解任务中,ILL使准确率最多降低67个百分点,而其平均人类可听度评分为1.33,接近干净音频的1.17;DRG在干净重采集后将平均受攻击准确率从28.5%提升至46.1%。这些发现明确了LALMs此前被忽视的安全风险,并为未来鲁棒音频理解研究提供了基础。
英文摘要
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active intervals and Frequency Confusion Transfer to construct a low-frequency waveform with continuous phase from corpus spectral variation. To mitigate this risk, we propose Distributional Requery Guard (DRG) to detect low-frequency distribution shifts and conditionally request a second recording for semantic recovery. Across six LALMs and multiple audio understanding tasks, ILL reduces accuracy by up to 67 percentage points while receiving a mean human audibility rating of 1.33, close to 1.17 for clean audio; DRG raises mean attacked accuracy from 28.5\% to 46.1\% after clean reacquisition. These findings identify a previously overlooked safety risk for LALMs and provide a foundation for future research on robust audio understanding.