发表机构
Hawkes Institute, University College London; AINOSTICS Ltd.(伦敦大学学院霍克斯研究所; AINOSTICS 有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过医学图像对齐评估任务测试前沿多模态模型的通用视觉推理能力,发现GPT-6等新模型性能优异,而本地微调模型迁移性较差,表明该任务可作为通用视觉推理的有效测试平台。
AI 中文摘要
前沿多模态大语言模型(MLLMs)正日益被定位为通用视觉推理器,作为追求人工通用智能的一部分。这种通用性的一个关键测试是,它们能否执行人类可以仅凭视觉证据和任务指令可靠做出的新颖视觉判断,而无需针对特定任务进行参数优化。我们通过医学图像对齐评估这一任务来研究这个问题,其目标是确定两幅图像之间是否存在解剖学对应关系。人类对图像对齐的视觉评估仍然是金标准且最常见的方法;然而,它需要训练有素的操作人员,并且对于大规模数据集而言难以扩展。我们在两个示例性医学图像对齐任务上评估了最近几代MLLMs,同时变化提示策略和图像呈现方法。我们与本地微调的MLLM和特定任务的CNN进行比较,以考察前沿通用模型与需要特定任务优化但可在本地使用的较小模型之间的权衡。我们表明,我们正在达到一个拐点,前沿MLLMs现在可以对医学图像对齐进行有效的视觉评估。仅几个月前发布的模型泛化能力差,在某些设置下,其表现仅略高于随机水平,而GPT-6在几乎所有测试场景中均达到超过85%的准确率。微调的本地模型在其训练的任务上可以匹配或超过前沿模型的性能,但对未见过的设置的迁移效果显著较差。这些发现将医学图像对齐确定为通用视觉推理的有用测试平台,并表明前沿多模态模型正开始获得能够支持跨异构医学影像流程的通用质量控制机制的能力。
英文摘要
Frontier multimodal large language models (MLLMs) are increasingly positioned as general purpose visual reasoners as part of the quest for artificial general intelligence. A key test of this generality is whether they can perform novel visual judgments that humans can make reliably from visual evidence and task instructions, without task-specific parameter optimisation. We investigate this question through the task of medical image alignment assessment, where the goal is to establish whether there is anatomical correspondence between two images. Human visual assessment of image alignment is still the gold standard and most common approach; however, it requires trained operators and is impractical to scale for large datasets. We evaluate recent generations of MLLMs on two exemplar medical image alignment tasks, varying both prompting strategies and image-presentation methods. We compare against a locally fine-tuned MLLM and a task-specific CNN to examine the trade-off between frontier general purpose models and smaller models that require specific task optimisation but can be used locally. We show that are reaching an inflection point, where frontier MLLMs can now perform effective visual assessment of medical image alignment. Models released only a few months ago generalise poorly and, in some settings, perform barely above chance, whereas GPT-6 achieves over 85% across almost all scenarios tested. Fine-tuned local models can match or exceed frontier-model performance on the tasks on which they are trained, but transfer substantially less effectively to unseen settings. These findings identify medical image alignment as a useful test bed for generalist visual reasoning and suggest that frontier multimodal models are beginning to acquire capabilities that could support a common quality-control mechanism across heterogeneous medical-imaging pipelines.