AI 中文总结
针对视频-LLM问答中引用幻觉问题,提出自验证流水线,用小型自然语言推理模型作稳定验证器,在对抗性错误前提问题上检测率达79%,并发布相关代码。
AI 中文摘要
基于视觉-语言模型构建的视频问答系统,即便引用的帧不支持相关内容,也常生成带有时间戳的高置信度断言。这种欺骗性幻觉源于时间戳虽暗示了 grounding 但无法保证正确性,这会提升用户信任度却无法提高准确率。本文提出了一个闭环流水线:检索增强语言模型生成带有每条断言时间戳引用的答案,且在向用户展示前,会对每个被引用的帧进行独立重新检查。我们将其与普通基线对比,并对三种验证设计进行 ablation 实验,在 Apple Silicon(MLX)和 Google Colab(HF Transformers、CUDA)上进行评估。直接询问视觉模型某帧是否支持某断言因谄媚性完全失败(40条断言上的检测率为0%);仅重新生成描述并结合通用 LLM 评判的方法结果不稳定,根据提示措辞的不同,标记率在0%到100%之间波动;用小型自然语言推理模型替代该评判器后,得到了稳定、可解释的验证器,在对抗性错误前提问题上能检测到79%的虚构断言,同时不影响真实断言。我们发布了完整流水线、评估工具,以及适用于 Apple Silicon 和 Colab 的实现代码,代码可在该 https URL 获取。
英文摘要
Video question answering systems built on vision-language models often produce timestamped claims with high confidence even when unsupported by the cited frame. This deceptive hallucination arises because timestamps imply grounding without ensuring correctness, increasing user trust but not accuracy. We introduce a pipeline that closes this loop. A retrieval-augmented language model drafts answers with per-claim timestamp citations, and each cited frame is independently re-examined before being shown to the user. We compare against a plain baseline and ablate three verification designs, evaluated on both Apple Silicon (MLX) and Google Colab (HF Transformers, CUDA). Directly asking the vision model whether a frame supports a claim fails completely (0% catch rate on 40 claims) due to sycophancy. Blind re-captioning plus a general LLM judge improves results but is unstable, oscillating between 0% and 100% flagged depending on prompt phrasing. Replacing that judge with a small natural language inference model yields a stable, interpretable verifier that catches 79% of fabricated claims on adversarial false-premise questions while leaving true claims untouched. We release the full pipeline, evaluation harness, and implementations for both Apple Silicon and Colab. Code is available at https://github.com/yogesh-iitj/grounded-video-qa.