arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10244cs.CL

双令牌特征与小-大模型集成用于VLM幻觉检测

Two-Token Features and Small-Large Ensembles for VLM Hallucination Detection

  • IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

Eli Schwartz

中文总结 AI 辅助

本研究提出一种结合小型微调VLM(4B参数)与大型零样本VLM(400B参数)的集成方法,利用双令牌特征和OCR进行字符级幻觉检测,在SHROOM-Visions 2026任务中取得良好排名。

中文摘要 AI 辅助

我们展示了针对SHROOM-Visions 2026共享任务(字符级VLM幻觉检测)所设计的系统。一个参数量为40亿的小型VLM被微调为逐令牌分类器,从其自身的隐藏状态中读取双令牌特征,并在预测时与一个约4000亿参数的零样本VLM评判器进行集成。两个组件均使用图像中可见文本的现成OCR结果。我们利用大型模型生成的合成幻觉数据作为集成多样性的来源,并通过验证集来选择特征层、训练数据和OCR接地方式。我们的官方提交在隐藏测试集上达到了平均Cor 0.487 / Cor-lbl 0.387,在任务的主要Cor-lbl指标上,分别位列英文第6/28、法文第6/21、意大利文第8/21和中文第7/22。

英文摘要

We present our system for the SHROOM-Visions 2026 shared task on character-level VLM hallucination detection. A small ($4$B-parameter) VLM is fine-tuned as a per-token classifier reading a two-token feature from its own hidden states, and is ensembled with a $\sim$400B zero-shot VLM judge at prediction time. Both components see off-the-shelf OCR of any visible in-image text. We use synthetic hallucination data generated by the large model as a source of ensemble diversity, and use validation to select feature layer, training data and OCR grounding. Our official entry reaches mean Cor $0.487$ / Cor-lbl $0.387$ on the hidden test set, placing $6$th/$28$ (EN), $6$th/$21$ (FR), $8$th/$21$ (IT) and $7$th/$22$ (ZH) on the task's primary Cor-lbl metric.

↑