发表机构
Beijing University of Posts and Telecommunications; Nanyang Technological University(北京邮电大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OmniConfess通过令牌级通道干预生成证据坦白,无需训练即可缓解全模态大模型的幻觉,并在OmniHalluBench基准上验证了有效性。
AI 中文摘要
全模态大语言模型(OmniLLMs)统一了文本、图像、音频和视频,但当生成依赖于错误的证据时会产生幻觉。现有的推理时方法可以减少幻觉,但很少揭示哪种证据维持了生成的承诺。我们提出了OmniConfess,一种无需训练的全模态幻觉缓解方法。它固定一个候选响应,并在受控的逐通道证据干预下以令牌分辨率重新评分,产生结构化的逐令牌-通道坦白,揭示响应的证据依赖性。OmniConfess利用这种坦白来保留有根据的内容,并纠正由无关或矛盾证据驱动的承诺。为了评估OmniConfess,我们构建了OmniHalluBench,一个包含3,540个示例的基准,由六个数据集构建,涵盖文本、图像、音频和视频设置以及判断和自由形式生成。实验表明,OmniConfess在异构模态和任务设置中缓解了幻觉。我们的代码和基准可在此https URL公开获取。
英文摘要
Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated commitment. We introduce OmniConfess, a training-free method for mitigating omni-modal hallucinations. It fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, producing a structured token-by-channel confession that reveals the response's evidential dependence. OmniConfess uses this confession to preserve grounded content and correct commitments driven by irrelevant or contradictory evidence. To evaluate OmniConfess, we construct OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation. Experiments show that OmniConfess mitigates hallucinations across heterogeneous modality and task settings. Our code and benchmark are publicly available at https://github.com/RongHuiQiang/OmniConfess.