arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11244cs.CLcs.CV

OmniHallu:多模态大语言模型中跨模态理解与生成的统一幻觉检测

OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

  • The University of Manchester(曼彻斯特大学)
  • The University of Hong Kong(香港大学)
  • Huazhong University of Science and Technology(华中科技大学)
  • Shanghai Academy of Educational Sciences(上海市教育科学研究院)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Jianjiang Yang, Peihang Li, Shanqing Xu, Mengchen Qian, Lu Zhang, Meng Luo

AI总结:

提出OmniHallu统一幻觉检测框架及含万样本的基准,覆盖六种跨模态任务,通过多智能体架构和偏好优化验证器减少专家调用66%,揭示模态相关性能梯度。

AI中文摘要:

尽管多模态大语言模型(MLLMs)在各种任务中取得了显著进展,但它们仍存在幻觉问题,即生成的输出与输入语义相矛盾或歪曲输入语义。现有研究通常仅在单一模态或任务类型内处理幻觉检测,限制了泛化能力。我们提出了OmniHallu,一个统一的幻觉检测框架,涵盖图像、视频和音频模态的理解与生成任务。我们贡献了OmniHallu-Bench,一个包含10,000个样本的基准,具有声明级人工标注,覆盖六种跨模态任务:图像到文本(I2T)、视频到文本(V2T)、音频到文本(A2T)、文本到图像(T2I)、文本到视频(T2V)和文本到音频(T2A)。我们的多智能体架构将模型输出分解为原子声明,通过模态特定专家进行验证,并通过结构化推理聚合证据。我们进一步提出了一种偏好优化的可训练验证器,近似多智能体决策边界,将专家调用减少66%,且性能损失极小。大量实验揭示了一致的模态相关性能梯度,并为跨模态幻觉模式提供了细粒度见解。

英文摘要:

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse tasks, they suffer from hallucinations where generated outputs contradict or misrepresent input semantics. Existing research typically addresses hallucination detection within a single modality or task type, limiting generalizability. We introduce OmniHallu, a unified hallucination detection framework spanning both comprehension and generation tasks across image, video, and audio modalities. We contribute OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations covering six cross-modal tasks: image-to-text (I2T), video-to-text (V2T), audio-to-text (A2T), text-to-image (T2I), text-to-video (T2V), and text-to-audio (T2A). Our multi-agent architecture decomposes model outputs into atomic claims, verifies them through modality-specific experts, and aggregates evidence via structured reasoning. We further propose a preference-optimized trainable verifier that approximates the multi-agent decision boundary, reducing expert calls by 66% with minimal performance loss. Extensive experiments reveal a consistent modality-dependent performance gradient and provide fine-grained insights into cross-modal hallucination patterns.

补充信息

↑