发表机构
Indian Institute of Science Education and Research; Vellore Institute of Technology(印度科学教育与研究学院; 韦洛尔理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对严重声学噪声下音频-文本基础模型的失效问题,提出无训练无数据源的TTA框架PRISM,在UrbanSound8K数据集上取得显著性能提升,还解决了复音陷阱问题。
AI 中文摘要
音频-文本基础模型(ATMs)在严重声学噪声下会出现灾难性失效,而现有的适应策略要么依赖基于梯度的测试时适应(TTA),该策略会强化噪声而非信号;要么依赖提示调优,这需要推理时无法获取的特权噪声注释。我们提出PRISM(Prototype-Rectified Iterative Self-supervised Manifold Denoising,原型修正迭代自监督流形去噪),这是一种无训练、无数据源的TTA框架,基于仿射噪声假设:严重声学噪声会在多模态潜在空间中引发低秩仿射偏移,超过90%的失真能量集中在前60个主成分中。PRISM利用冻结的文本原型作为几何锚点,通过仿射偏差回归将三种闭式几何修正编译为单个静态投影矩阵,从未标记的目标批次中估计并反转该失真。在推理时,适应过程简化为一次矩阵-向量乘法,耗时0.0009毫秒,比基于梯度的TTA快得多,且无需额外训练。在UrbanSound8K数据集上,PRISM较零样本基线提升12.94个百分点,超过特权辅助TTA基线9.41个百分点,尽管从未观察到其特权增强噪声提示。我们进一步识别出复音陷阱——宽带类子空间收缩的一种原则性失效模式,并通过置信度感知回归(CAR)解决该问题,为受影响最严重的类恢复了多达8.16个百分点的性能。
英文摘要
Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We address these failures with PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis: severe acoustic noise induces a low-rank affine shift in the multimodal latent space, with more than 90% of distortion energy confined to the leading 60 principal components. PRISM estimates and reverses this distortion from an unlabeled target batch using frozen text prototypes as geometric anchors via three closed-form geometric corrections compiled into a single static projection matrix by Affine Bias Regression. At inference, adaptation reduces to one matrix-vector multiplication in 0.0009 ms, making it substantially faster than gradient-based TTA while requiring no additional training. On UrbanSound8K, PRISM improves over the zero-shot baseline by 12.94 percentage points and surpasses an oracle-assisted TTA baseline by 9.41 percentage points, despite never observing its privileged augmented noise prompts. We further identify the Polyphonic Trap, a principled failure mode of subspace deflation for broadband classes, and resolve it via Confidence-Aware Regression (CAR), recovering up to 8.16 percentage points for the worst-affected class.
CommentsAccepted as a full paper at ACM CIKM 2026