arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

科学图像质量评估:基于多模态检索增强生成

Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation

Yinuo Zhang, Bingshuo Liu, Zhiying Tu, Dianhui Chu, Qingbin Liu, Xi Chen, Jiang Bian, Xiaoyan Yu, Dianbo Sui

arXiv 2609.19634首次发表:更新:

发表机构

Harbin Institute of Technology; Tencent; Nanyang Technological University(哈尔滨工业大学; 腾讯; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于多模态检索增强生成的科学图像质量评估框架,同时处理理解与评分任务,通过多路检索融合提升评估能力,并在ICME 2026挑战赛中获SIQA-U赛道第一。

AI 中文摘要

本文提出了一种用于科学图像质量评估的检索增强生成(RAG)框架,旨在同时解决SIQA挑战中的理解赛道(SIQA-U)和评分赛道(SIQA-S)。我们构建了一个整合文本语义与细粒度视觉特征的多模态索引,并开发了多路检索与融合机制,为大型语言模型提供高度相关的参考案例,从而增强其评估复杂科学图像的能力。实验结果表明,所提框架有效对齐了人类专家的判断标准。最终,我们的方法在ICME 2026大挑战赛的SIQA-U赛道中荣获第一名。

英文摘要

This paper proposes a Retrieval-Augmented Generation (RAG) framework for scientific image quality assessment, designed to simultaneously address both the understanding track (SIQA-U) and the scoring track (SIQA-S) of the SIQA challenge. We construct a multimodal index that integrates textual semantics with fine-grained visual features, and develop a multi-route retrieval and fusion mechanism to provide large language models with highly relevant reference cases, thereby enhancing their capability to evaluate complex scientific images. Experimental results demonstrate that the proposed framework effectively aligns with the judgment criteria of human experts. Ultimately, our method achieves 1st place in the SIQA-U track of the SIQA challenge at the ICME 2026 Grand Challenges.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑