arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36440cs.CV

DARE 缓解幻觉:双路径自回归感知编辑

DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing

Jae-Ho Lee, Jeong-Eun Lee, Gyeong-Moon Park

首次发表
浏览论文内容

中文总结 AI 辅助

针对大型视觉语言模型中的对象幻觉问题,提出融合文本与图像双路径及自回归感知信号的DARE编辑框架,有效减少幻觉并保持多模态性能。

中文摘要 AI 辅助

大型视觉语言模型(LVLMs)近年来在多模态任务上取得了显著进展,但对象幻觉仍然是一个持续存在的挑战,即模型生成的描述与视觉输入不一致。近期工作通过免训练的表示编辑来缓解幻觉,通常利用教师强制(TF)对比在幻觉响应与真实响应之间构建与幻觉相关的方向。然而,LVLMs 在生成过程中通过自回归(AR)解码运行,这引发了一个问题:基于 TF 的分析是否完全反映了导致幻觉输出的生成动态。本文分析了基于 TF 的编辑与 AR 生成行为之间的关系,发现仅基于 TF 的编辑可能不足以捕捉与幻觉相关的解码动态和多模态交互。为解决这一局限,我们提出 DARE(双路径自回归感知编辑),一种混合幻觉编辑框架,集成了两条互补的对比路径:文本对比和图像对比,以及自回归感知的表示信号。具体而言,DARE 从以下来源构建幻觉编辑方向:(1)基于 TF 的文本对比,(2)解码过程中自回归感知的表示转换,以及(3)配对图像之间的受控视觉差异。在多个 LVLM 幻觉基准上的大量实验表明,DARE 持续减少对象幻觉,同时保持多模态感知能力和推理效率。我们的实现代码可在以下网址获取。

英文摘要

Large vision-language models (LVLMs) have recently achieved remarkable progress across multimodal tasks, yet object hallucination remains a persistent challenge where models generate descriptions inconsistent with the visual input. Recent work mitigates hallucinations through training-free representation editing, typically by constructing hallucination-related directions from teacher-forcing (TF) contrasts between hallucinated and truthful responses. However, LVLMs operate through autoregressive (AR) decoding during generation, raising the question of whether TF-based analysis fully reflects the generation dynamics that lead to hallucinated outputs. In this paper, we analyze the relationship between TF-based editing and AR generation behavior and find that TF-based editing alone may be insufficient to capture both decoding dynamics and multimodal interactions associated with hallucinations. To address this limitation, we propose DARE (Dual-path Auto-Regressive-aware Editing), a hybrid hallucination editing framework that integrates two complementary contrast pathways: textual contrasts and image contrasts, together with autoregressive-aware representation signals. Specifically, DARE constructs hallucination editing directions from (1) TF-based textual contrasts, (2) AR-aware representation transitions during decoding, and (3) controlled visual differences between paired images. Extensive experiments on multiple LVLM hallucination benchmarks demonstrate that DARE consistently reduces object hallucinations while preserving multimodal perception capability and inference efficiency. Our implementation code is available at https://github.com/KU-VGI/DARE.

发表机构

  • Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑