REVE:通过复用编码器状态实现大型音频语言模型的高效幻觉纠正
REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States
- Beijing Institute of Technology, Zhuhai, Guangdong, China(北京理工大学(珠海))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
REVE通过复用大型音频语言模型已计算的编码器状态,以轻量方式验证并纠正幻觉事件提及,在AudioSet上移除了92.9%的标签不支持提及,且延迟仅为CED-Base的约1/18。
AI中文摘要:
大型音频语言模型可能会提及输入中不存在的声学事件。单独的音频事件检测器可以验证这些提及,但这需要第二个音频编码器和一次单独的前向传播。我们提出了用于验证事件的重用编码器状态(REVE),这是一种轻量级方法,利用目标模型已经计算出的状态。一个读出器汇总跨音频帧的类别分数,而另一个读出器使用来自四个连续帧区间的池化状态。类别感知分数融合结合它们的输出来验证生成的事件提及,而无需再次对音频进行编码。在AudioSet上,在忠实提及召回约束下,REVE移除了92.9%的标签不支持提及。凭借更少的添加参数且无需第二次音频编码传递,REVE实现了与CED-Tiny和CED-Base相当的减少效果。其完整验证延迟约为CED-Base路径的1/18。在受控DESED混合物和不同目标模型架构上的结果进一步证实了编码器状态复用的有效性。
英文摘要:
Large audio-language models may mention acoustic events that are absent from the input. A separate audio event detector can verify these mentions, but doing so requires a second audio encoder and a separate forward pass. We propose Reused Encoder States for Verifying Events (REVE), a lightweight method that uses states already computed by the target model. One readout summarizes class scores across audio frames, while another uses pooled states from four consecutive frame intervals. Class-aware score fusion combines their outputs to verify generated event mentions without encoding the audio again. On AudioSet, REVE removes 92.9% of label-unsupported mentions under a faithful-mention recall constraint. With fewer added parameters and no second audio-encoding pass, REVE achieves a reduction comparable to those of CED-Tiny and CED-Base. Its complete verification latency is about 1/18 of the CED-Base path. Results on controlled DESED mixtures and different target-model architectures further confirm the effectiveness of encoder-state reuse.