AI 中文总结
提出VidHalLoc基准和VideoHALO工作流,统一评估视频幻觉检测器,发现专用检测器总体准确率仅34.63%,可靠性有限。
AI 中文摘要
视频语言模型和视频智能体可能产生与时空证据相冲突的幻觉。现有基准主要评估模型幻觉,而异构机制使得检测器的可靠性难以比较。我们引入了VidHalLoc,一个在统一诊断评估协议下评估幻觉检测方法的基准,使用2000个对抗性幻觉样本,涵盖视频问答和视频字幕任务,跨越本体和动态幻觉类别。为高效构建VidHalLoc,我们引入了VideoHALO,一个基于工程信息的多智能体工作流,将数据构建分解为四个可执行阶段,由记忆系统和通信协议支持。对十五种方法的评估显示,四个专用检测器的总体准确率峰值仅为34.63%,表明在视频幻觉类型上的可靠性有限。[数据集仓库:此https URL]
英文摘要
Video-language models and video agents can produce hallucinations that conflict with spatiotemporal evidence. Existing benchmarks mainly evaluate model hallucinations, and heterogeneous mechanisms make detector reliability difficult to compare. We introduce VidHalLoc, a benchmark that evaluates hallucination detection methods under a unified diagnostic evaluation protocol using 2,000 adversarial hallucination samples across Video Question Answering and Video Captioning tasks, spanning Ontology and Dynamic hallucination categories. To construct VidHalLoc efficiently, we introduce VideoHALO, a Harness Engineering-informed multi-agent workflow that decomposes data construction into four executable stages supported by a memory system and a communication protocol. Evaluation of fifteen methods reveals that the four dedicated detectors peak at an Overall accuracy of only 34.63%, indicating limited reliability across video hallucination types [Dataset Repository: https://huggingface.co/datasets/wesfggfd/VidHalLoc].
Comments29 pages, including appendices