发表机构
Hokkaido University; Neurogica Inc.(北海道大学; Neurogica公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对冻结多模态大模型在对抗性视觉定位基准上易出错的问题,提出无需训练的测试时融合方法FORUM,利用模型间一致性及几何规则选择目标框,在Ref-Adv-s上平均准确率相对超过397B模型5%。
AI 中文摘要
冻结的多模态大语言模型(MLLMs)现在可以通过一次提示调用解决标准的指代表达理解任务,然而在面对具有同类干扰物和否定词的对抗性基准测试时,即使最大的模型也会自信地出错,而重新采样会重复同样的错误。由不同数据和架构构建的模型很少会陷入相同的混淆因素,因此它们之间的一致性是一种无需标签的强信号,指示正确的目标。我们提出FORUM,一种无需训练的测试时融合方法,用于融合冻结的MLLMs,由两条固定的几何规则指导:基于一致性的选择保留得到最多不同模型支持的候选区域,而medoid定位返回一个实际的成员框而非坐标平均值,因此一个松散的预测不会使答案偏移。融合三个开源的MLLMs,FORUM在对抗性Ref-Adv-s基准上的平均准确率相对超过397B参数的已发表参考模型5%,并超过简单平均集成15%。这些增益可迁移到标准的RefCOCO+基准,且在没有主导成员的平衡阵容中,仍超过397B模型5%。
英文摘要
Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial benchmarks with same-category distractors and negation, even the largest models are confidently wrong, and resampling repeats the error. Models built from different data and architectures rarely fall for the same confounder, so their agreement is a strong label-free signal of the correct target. We present FORUM, a training-free test-time fusion of frozen MLLMs guided by two fixed geometric rules: agreement-based selection keeps the region supported by the most distinct models, and medoid localization returns an actual member box instead of a coordinate average, so one loose prediction cannot shift the answer. Fusing three open MLLMs, FORUM surpasses the 397B-parameter published reference by a relative 5% in mean accuracy on the adversarial Ref-Adv-s benchmark, and a plain averaging ensemble by 15%. The gains transfer to standard RefCOCO+, and a balanced lineup with no dominant member still surpasses the 397B model by 5%.
CommentsAccepted by ACCV 2026