发表机构
Center for Computational and Data Sciences Lab, Independent University Bangladesh; Independent Author(孟加拉国独立大学计算与数据科学中心实验室; 独立作者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HALDETECT 系统采用答案优先对比接地与 QLoRA 微调,在 ImageEval 2026 幻觉检测任务中取得第三名,验证了适配方法优于提示,并揭示了数据扩展的种子敏感性。
AI 中文摘要
大型多模态模型倾向于流畅地产生视觉细节的幻觉,这限制了它们在细粒度解释方面的部署。我们提出了 HALDETECT,这是我们针对 ImageEval 2026 英语幻觉检测赛道(任务 1b)的系统,在该任务中,系统必须从一张图像和三个文化上合理的陈述中,识别出唯一有视觉依据的那一个。我们将该问题构建为一个对比决策,在解释之前先输出答案,并围绕颜色/纹理、形状/形式以及上下文来组织推理。我们提交的最佳适配器使用 4 位 QLoRA 对 Qwen2.5-VL-7B-Instruct 进行微调,同时冻结视觉编码器,并在 1,000 项测试集上达到了对比不稳定性(CI)0.035;我们在八支队伍中排名第三。开发实验表明,答案顺序可能比模型规模更重要,并且适配方法优于仅使用提示。对已发布黄金标签的回顾性配对分析证实了 QLoRA 相对于最佳提示的优势,但未证实开发测试集选择的适配器与最佳测试适配器之间的微小差距,并且对全部四种训练规模重新设置种子表明,表观的数据扩展曲线在种子变化后不再成立。剩余的 35 个残差错误涉及文化上合理的功能、材质和识别区分;朴素的适配器投票没有帮助。
英文摘要
Large multimodal models tend to hallucinate visual detail fluently, which limits their deployment for fine-grained interpretation. We present HALDETECT, our system for the English hallucination-detection track (Task 1b) of ImageEval 2026, in which a system must identify, from an image and three culturally plausible statements, the single visually grounded one. We frame the item as one contrastive decision, emit the answer before its explanation, and structure reasoning around colour/texture, shape/form, and context. Our best submitted adapter fine-tunes Qwen2.5-VL-7B-Instruct with 4-bit QLoRA while freezing the vision encoder and reaches Contrastive Instability (CI) 0.035 on the 1,000-item test set; we placed third of eight teams. Development experiments show that answer order can matter more than model scale and that adaptation beats prompting alone. Retrospective paired analysis of the released gold labels confirms the QLoRA gain over the best prompt but not the small gap between the devtest-selected and best-test adapters, and reseeding all four training sizes shows that the apparent data-scaling curve does not survive a seed change. The 35 residual errors are culturally plausible function, material, and recognition distinctions; naive adapter voting does not help.
Comments10 pages, 3 figures, 10 tables (including appendices). System description paper for Task 1b (English) of ImageEval 2026 Shared Tasks (Fourth Arabic Natural Language Processing Conference), to appear in the Shared Tasks proceedings