Comments17 pages, 14 figures, accepted to Computer Vision and Pattern Recognition Conference (CVPR) Workshops 2026. 5th MMFM Workshop: What is Next in Multimodal Foundation Models?
Journal refIn Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7415-7424) 2026
CommentsThis article presents only the preliminary research results, which are not yet complete and lack necessary supplementary experiments. The author has decided to withdraw it to improve the research work, and will submit a more complete version in the future
Jiageng Wen, Shengjie Zhao, Bing Li, Jiafeng Huang, Kenan Ye, Hao Deng
机构
*
Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学智能自主系统上海研究院)
;
School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
;
School of Mechatronic Engineering and Automation, Shanghai University(上海大学机械电子工程与自动化学院)
CommentsAccepted to ICLR 2026. This arXiv version includes an additional appendix (Appendix 15) containing further philosophical discussion not included in the official ICLR peer-reviewed version
Yuting Wan, Liguo Sun, Jiuwu Hao, Zao Zhang, Pin LV
机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
MixRI: Mixing Features of Reference Images for Novel Object Pose Estimation
MixRI:用于新颖物体姿态估计的参考图像特征混合
Xinhang Liu, Jiawei Shi, Zheng Dang, Yuchao Dai
机构
*
School of Electronics and Information, Northwestern Polytechnical University(电子与信息学院,西北工业大学)
;
Shaanxi Key Laboratory of Information Acquisition and Processing(陕西省信息获取与处理重点实验室)
;
CVLab, EPFL, Switzerland(EPFL瑞士计算机视觉实验室)
Balancing Accuracy and Efficiency: CNN Fusion Models for Diabetic Retinopathy Screening
在准确性和效率之间平衡:用于糖尿病视网膜病变筛查的CNN融合模型
Md Rafid Islam, Rafsan Jany, Akib Ahmed, Mohammad Ashrafuzzaman Khan
机构
*
Department of Electrical and Computer Engineering(电气与计算机工程系)
;
North South University(北南大学)
;
Digital Health Research Division(数字健康研究部)
;
Korea Institute of Oriental Medicine(韩国东方医学研究院)
;
Department of Computer Science(计算机科学系)
;
American International University--Bangladesh(美国国际大学-孟加拉国)
DAGLFNet: Deep Feature Attention Guided Global and Local Feature Fusion for Pseudo-Image Point Cloud Segmentation
DAGLFNet:基于伪图像的深度特征注意力引导的全局和局部特征融合用于伪图像点云分割
Chuang Chen, Yi Lin, Bo Wang, Jing Hu, Xi Wu, Wenyi Ge
机构
*
College of Computer Science, Chengdu University of Information Technology(成都信息科技大学计算机学院)
;
College of Computer Science, Sichuan University(四川大学计算机学院)
;
Publications Department, Optica Publishing Group(Optica出版社出版部)
;
Department of Electronic Journals, Optica Publishing Group(Optica出版社电子期刊部)
Depth-Supervised Fusion Network for Seamless-Free Image Stitching
Zhiying Jiang, Ruhao Yan, Zengxi Zhang, Bowei Zhang, Jinyuan Liu
机构
*
College of Information Science and Technology, Dalian Maritime University(大连海事大学信息科学与技术学院)
;
School of Software Technology, Dalian University of Technology(大连理工大学软件学院)
Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation
Jeonghyun Noh, Wangsu Jeon, Jinsun Park
机构
*
Department of Information Convergence Engineering, Pusan National University(convergence工程系,釜山国立大学)
;
School of Computer Engineering, Kyungnam University(计算机工程学院,庆尚大学)
;
School of Computer Science and Engineering, Pusan National University(计算机科学与工程学院,釜山国立大学)
机构
*
School of Physics and Technology(物理与技术学院)
;
Wuhan University(武汉大学)
;
Pratt School of Engineering(工程学院)
;
Duke University(杜克大学)
;
Faculty of Artificial Intelligence in Education(教育人工智能学院)
;
Central China Normal University(中部师范大学)
;
Huangpu Branch of Shanghai Ninth People’s Hospital(上海第九人民医院黄浦分院)