发表机构
Shaggar Institute of Technology; Trinity College Dublin(沙加尔理工学院; 都柏林圣三一学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SpanCalib-VLM混合双系统,结合多模态序列标注器与微调生成式VLM,经联合校准融合策略优化,在SHROOM-Visions任务上实现幻觉跨度的高效校准检测,公开了模型权重与代码。
AI 中文摘要
检测大型视觉语言模型(LVLMs)中的幻觉,既需要准确的跨度定位,也需要校准良好的置信度分数。微调后的生成式视觉语言模型(VLMs)在识别幻觉文本跨度方面表现出色,但存在过度自信和推理延迟高的问题;判别式序列标注器速度确定且校准效果更优,但跨度召回率偏保守。本文提出SpanCalib-VLM,这是用于SHROOM-Visions共享任务的混合双系统,将由XLM-RoBERTa-Large融合SigLIP视觉编码器(通过交叉注意力构成)的多模态序列标注器,与微调后的生成式VLM(Qwen3.5-4B-SHROOM-SFT)相结合;通过联合校准融合策略,用序列标注器的校准概率对生成式模型的候选跨度重新评分。在SHROOM-Visions英文评估拆分集上,该集成模型达到0.41的皮尔逊校准相关系数、0.39的总体交并比(IoU)、干净响应的IoU为0.91,总体检测准确率为70.7%,且已公开模型权重与代码。
英文摘要
Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-calibrated confidence scores. Fine-tuned generative VLMs excel at identifying hallucinated text spans but suffer from overconfidence and high inference latency. Discriminative sequence taggers offer deterministic speed and superior calibration but exhibit conservative span recall. We present SpanCalib-VLM, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with our fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT). Through a Union-Calibrated Fusion strategy, candidate spans from the generative model are re-scored with calibrated probabilities from the sequence tagger. On the SHROOM-Visions English evaluation split, our ensemble achieves a Pearson calibration correlation of 0.41 and an overall IoU of 0.39, with a clean-response IoU of 0.91} and overall detection accuracy of 70.7%. We make our model weights and code publicly available.
CommentsShroom-Visions