AI 中文总结
针对多模态大语言模型的长文本描述对象幻觉问题,本文提出双流跨锚校正方法,通过耦合感知流与认知流提升精度,在长文本场景下实现最优性能,且存在域条件性限制。
AI 中文摘要
多模态大语言模型中的对象幻觉问题,是当语言先验和语料共现偏差超过视觉证据时产生的,没有将单个对象提及与图像实际内容关联起来。大多数补救措施在解码时进行干预,无需训练,但在统一协议下,其优势仅限于短文本描述;在细节丰富的语料库上进行监督微调(SFT)可延长文本描述,但仍有超过40%的文本描述会提及不存在的对象。本文提出双流跨锚校正(DSCC),与解码后处理的工作不同,DSCC是首个在微调期间将对象级视觉锚点注入语言模型本身的方法:感知流通过双向对比目标将中间层的对象级隐藏状态与冻结的文本锚点对齐;认知流让更深层在每个生成步骤通过交叉注意力查询这些锚点;两阶段课程门将两者耦合,使证据检索成为每个自回归步骤的结构约束。在同一主干网络和同一评分协议下,实验涵盖长文本描述幻觉、对象存在判别和跨域泛化,以相同语料库和相同计划的普通SFT作为长度和密度匹配的对照,因此增益可逐层归因。DSCC是唯一达到长文本描述、低幻觉区域的方法:文本描述长度约为基线的1.9倍,每个对象提及的精度为88.19%,是密度独立准则下的最高值。消融实验揭示了协同作用:单独的感知流会降低精度,但若叠加在认知流上则会反转符号。本文不声称具有普遍优越性:三个域外基准产生可预测、可证伪的域条件性,该协同作用受限于锚点的语义域,在图表和视错觉上会失效。
英文摘要
Object hallucination in multimodal large language models arises when language priors and corpus co-occurrence bias outweigh the visual evidence, with nothing tying an object mention to the image. Most remedies intervene at decoding time, yet under a unified protocol their benefit is confined to short captions; supervised fine-tuning (SFT) on a detail-rich corpus lengthens captions, but over forty percent still name absent objects. This paper proposes Dual-Stream Cross-Anchor Correction (DSCC). Unlike work that post-processes decoding, DSCC injects object-level visual anchors into the language model itself during fine-tuning: a perception stream aligns object-level hidden states at an intermediate layer to frozen text anchors by a bidirectional contrastive objective; a cognition stream lets deeper layers query those anchors by cross-attention at every generation step; and a two-stage curriculum gate couples them, making evidence retrieval a structural constraint on generation. Under one backbone and one scoring protocol, experiments span long-caption hallucination, object-existence discrimination and cross-domain generalisation, with vanilla SFT on the same corpus and schedule as a length- and density-matched control separating the data effect from the architectural gain. DSCC alone reaches the long-caption, low-hallucination region: captions roughly 1.9 times the baseline length at 88.19% precision per object mention, the highest under a density-independent criterion. Ablations expose a synergy: the perception stream alone degrades precision yet reverses sign when stacked on the cognition stream. No universal superiority is claimed: three out-of-domain benchmarks yield a predictable, falsifiable domain-conditionality, the synergy being bound to the anchors' semantic domain and breaking on charts and illusions.