以人类方式进行图像融合排序:用于红外-可见光融合评估的学习型成对偏好度量
Ranking Infrared-Visible Fusion the Way Humans Do: A Learned Pairwise Preference Measure
浏览论文内容
中文总结 AI 辅助
针对红外-可见光图像融合缺乏理想参考、成对比较成本高的问题,提出LPIFM模型,利用新偏好语料库训练,可准确复现人类成对判断,性能优于传统度量,还公开了相关数据集与代码。
中文摘要 AI 辅助
红外-可见光图像融合(IVIF)不存在理想的融合参考,因此融合算法通常通过标量客观度量进行排序,这些度量将信息传递、结构或源相似性的不同代理形式化。这些代理常常与最终重要的判断不一致:在相同源的情况下,人类更偏好两个融合结果中的哪一个?直接成对比较是相对主观评估的既定参考协议,但它的成本随算法数量呈二次增长,阻碍了其常规使用。我们提出学习型感知图像融合度量(LPIFM),这是一种源条件模型,将人类A/B/平局比较协议实现为可重复、可扩展的替代方案。LPIFM共同观察红外源、可见光源和两个融合候选,并预测候选A更好、候选B更好还是两者在感知上等价。监督来自一个新的密集偏好语料库,该语库涵盖了公共基准场景中大量融合方法之间的所有无序比较,采用盲法、随机化的两阶段协议并经专家裁决进行标注。在场景和方法泛化设置中,LPIFM紧密跟踪人类成对决策,并复现源自人类标签的感知平局感知Bradley-Terry排序;在完整方法池上,它在成对准确率和排序相关性方面大幅优于最强的传统度量。我们发布带注释的偏好数据集,以及LPIFM模型权重、源代码和评估代码,以支持偏好对齐的IVIF评估。LPIFM为大规模人类对齐的方法比较和排序提供了实用工具。
英文摘要
Human pairwise comparison provides a direct basis for perceptual infrared-visible image fusion assessment, but dense annotation becomes costly as method pools grow. We present the Learned Perceptual Image Fusion Measure (LPIFM), among the earliest learned fusion assessors trained directly on dense human A/B/Tie comparisons. LPIFM jointly examines both source images and both fused candidates, combining a shared hierarchical encoder, triadic interaction, and a tie-aware objective to predict comparative preference and perceptual indifference. We construct and publicly release all 6,300 unordered comparisons among 25 methods on 21 VIFB scenes, collected through blinded, randomized annotation and expert adjudication. Across four VIFB evaluation settings, LPIFM achieves 79.2-84.0% agreement with human pairwise judgments and Spearman correlations of 0.941-0.977 with human-derived method rankings. On full method pools, accuracy exceeds the strongest of 19 conventional metrics by 16.3-21.1 pp. Consistency diagnostics show 99.98-100% candidate-swap agreement and no observed decisive preference cycles. External experiments on EVAFusion further demonstrate rapid adaptation to a different fusion-evaluation preference protocol. After only three epochs of fine-tuning, LPIFM surpasses all 19 conventional metrics in accuracy, macro-F1, and ranking correlation. LPIFM provides a scalable instrument for human-aligned fusion assessment, with the preference corpus, model weights, and code publicly available.