双令牌特征与小-大模型集成用于VLM幻觉检测
Two-Token Features and Small-Large Ensembles for VLM Hallucination Detection
- IBM Research(IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究提出一种结合小型微调VLM(4B参数)与大型零样本VLM(400B参数)的集成方法,利用双令牌特征和OCR进行字符级幻觉检测,在SHROOM-Visions 2026任务中取得良好排名。
中文摘要 AI 辅助
我们展示了针对SHROOM-Visions 2026共享任务(字符级VLM幻觉检测)所设计的系统。一个参数量为40亿的小型VLM被微调为逐令牌分类器,从其自身的隐藏状态中读取双令牌特征,并在预测时与一个约4000亿参数的零样本VLM评判器进行集成。两个组件均使用图像中可见文本的现成OCR结果。我们利用大型模型生成的合成幻觉数据作为集成多样性的来源,并通过验证集来选择特征层、训练数据和OCR接地方式。我们的官方提交在隐藏测试集上达到了平均Cor 0.487 / Cor-lbl 0.387,在任务的主要Cor-lbl指标上,分别位列英文第6/28、法文第6/21、意大利文第8/21和中文第7/22。
英文摘要
We present our system for the SHROOM-Visions 2026 shared task on character-level VLM hallucination detection. A small ($4$B-parameter) VLM is fine-tuned as a per-token classifier reading a two-token feature from its own hidden states, and is ensembled with a $\sim$400B zero-shot VLM judge at prediction time. Both components see off-the-shelf OCR of any visible in-image text. We use synthetic hallucination data generated by the large model as a source of ensemble diversity, and use validation to select feature layer, training data and OCR grounding. Our official entry reaches mean Cor $0.487$ / Cor-lbl $0.387$ on the hidden test set, placing $6$th/$28$ (EN), $6$th/$21$ (FR), $8$th/$21$ (IT) and $7$th/$22$ (ZH) on the task's primary Cor-lbl metric.