VADER:用于视频大语言模型幻觉缓解的自适应去偏方法
VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models
AI总结:
本研究针对视频大语言模型的幻觉问题,提出无训练自适应去偏框架VADER,通过视觉焦点重分配与选择性证据擦除模块结合对比解码,在LLaVA-Video-7B等模型的EventHallusion任务上取得显著性能提升。
AI中文摘要:
大型视觉语言模型(LVLMs)在开放式视频理解任务中展现出强大性能,但仍易生成缺乏视频证据支撑的流畅响应。现有无训练方法通常采用全局固定的视觉干预,或通过输入扰动构建对比分支;前者无法适配依赖视频的融合路径,后者可通过跨帧冗余得到补偿。为此,我们提出无训练框架视频自适应去偏方法(VADER),包含两个互补模块:视觉焦点重分配(VFR)为每个视频-问题输入自动实例化干预策略,诊断逐层视觉到文本的证据流,确定干预位置并推导预softmax注意力从系统token到视频token块的重分配强度;选择性证据擦除(SEE)独立掩码每一帧中高重要性视觉token,构建难以通过相邻帧补偿的先验偏差分支,对比解码随后降低选择性证据擦除后仍保持自信的预测权重。在多个视频大语言模型上,VADER在事件级接地和时间一致性任务中取得显著提升,在LLaVA-Video-7B模型上,其在EventHallusion数据集上达到72.60%的准确率。
英文摘要:
Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses unsupported by video evidence. Existing training-free methods typically apply a globally fixed visual intervention or construct a contrastive branch through input perturbation. The former cannot accommodate video-dependent fusion paths, while the latter can be compensated by cross-frame redundancy. We therefore propose Video-Adaptive Debiasing via Evidence Reweighting (VADER), a training-free framework with two complementary modules. Visual Focus Reallocation (VFR) automatically instantiates an intervention policy for each video-question input: it diagnoses layer-wise visual-to-text evidence flow, determines where to intervene, and derives how strongly to reallocate pre-softmax attention from system-token to video-token blocks. Selective Evidence Erasure (SEE) independently masks high-importance visual tokens in every frame, constructing a prior-biased branch that is difficult to compensate through neighboring frames. Contrastive decoding then down-weights predictions that remain confident after selective evidence erasure. Across multiple VideoLLMs, VADER yields substantial improvements on event-level grounding and temporal consistency; on LLaVA-Video-7B, it reaches 72.60% accuracy on EventHallusion.