发表机构
Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃尔兰根-纽伦堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FAU提交的多模态推理系统,未做任务特定模型训练,在2026年ImageCLEF竞赛中,获Visual MCQ第三名、Visual OpenQA第一名,凸显推理工程的实用价值。
AI 中文摘要
我们提出了用于视觉多项选择题问答(Visual MCQ)和视觉开放式问答(Visual OpenQA)子任务的2026年ImageCLEF多模态推理系统。该挑战要求对包含密集文本、图表、图形、表格、公式和单位的多语种教育及科学图像进行可靠推理,同时需遵守严格的答案格式要求。我们的核心发现是,鲁棒的输出控制与模型选择同等重要。对于Visual MCQ,我们用从视觉语言模型(VLM)的logits直接进行候选标签评分替代脆弱的自由格式生成,随后通过评分融合与投票组合互补运行结果。对于Visual OpenQA,我们采用图像增强、简洁最终答案提示、确定性解码以及针对性后处理以去除推理痕迹和格式伪影。未进行特定任务的模型训练,我们的官方提交在Visual MCQ中以0.7108的准确率获得第三名,在Visual OpenQA中以0.6488的COMET值、0.1391的BLEU值、0.2762的ROUGE L值和0.2383的METEOR值获得第一名。这些结果凸显了推理工程的实用价值:细致的评分、集成、提示和清理可将强大的VLM转化为可靠的竞赛系统。
英文摘要
We present our ImageCLEF 2026 Multimodal Reasoning system for the Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question Answering (Visual OpenQA) subtasks. The challenge requires reliable reasoning over multilingual educational and scientific images with dense text, diagrams, charts, tables, formulas, and units, while enforcing strict answer formats. Our central finding is that robust output control is as important as model choice. For Visual MCQ, we replace fragile free-form generation with direct candidate label scoring from vision-language model logits, then combine complementary runs through score fusion and voting. For Visual OpenQA, we use image enhancement, concise final answer prompting, deterministic decoding, and targeted post-processing to remove reasoning traces and formatting artifacts. Without task-specific model training, our official submissions achieved third place in Visual MCQ with 0.7108 accuracy and first place in Visual OpenQA with 0.6488 COMET, 0.1391 BLEU, 0.2762 ROUGE L, and 0.2383 METEOR. The results highlight the practical value of inference engineering: careful scoring, ensembling, prompting, and cleanup can turn strong VLMs into reliable competition systems.
Comments16 pages, 3 figures, 7 tables. CLEF 2026 Working Notes, ImageCLEF 2026 Multimodal Reasoning Task