MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
MCERF:通过增强检索推进工程文档的多模态大语言模型评估
Kiarash Naghavi Khanghah, Hoang Anh Nguyen, Anna C. Doris, Amir Mohammad Vahedi, Daniele Grandi, Faez Ahmed, Hongyi Xu
机构
*
School of Mechanical, Aerospace, and Manufacturing Engineering, University of Connecticut, Storrs, CT 06269(机械、航空航天与制造工程学院,康涅狄格大学,斯托尔斯,CT 06269)
;
Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA(机械工程系,麻省理工学院,剑桥,MA 02139,美国)
Comments9 pages. v2: results updated to July 2026 leaderboard (17 models). Accepted at the 2nd Workshop on Knowledge-Intensive Multimodal Reasoning (KnowledgeMR) at CVPR 2026 (non-archival), under the former title "PDFParse: A Benchmark for Grounded Multimodal Reasoning over Professional PDF Documents". Dataset: https://huggingface.co/datasets/surgeai/GDP.pdf ; Code: https://github.com/surge-ai/gdp-pdf
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
MELLA:弥合低资源语言多模态大语言模型的语言能力与文化根基
Yufei Gao, Jiaying Fei, Nuo Chen, Ruirui Chen, Guohang Yan, Yunshi Lan, Botian Shi
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
East China Normal University(东华大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Institute of High Performance Computing, A*STAR(高性能计算研究所,A*STAR)
Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation
通过自我场景增强在多模态大语言模型中强化自我中心空间感知
Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng, Zidong Cao, Lutao Jiang, Zixin Zhang, Huiyu Zhou, Xuming Hu
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Guangxi Zhuang Autonomous Region Information Center(广西壮族自治区信息中心)
;
The Hong Kong University of Science and Technology(香港科技大学)
Million-scale multimodal pollen microscopy with expert-guided foundation models
百万级多模态花粉显微镜图像与专家引导的基础模型
András Biricz, Björn Gedda, Donát Magyar, Antonio Spanu, János Fillinger, Péter Pollner, István Csabai
机构
*
Department of Physics of Complex Systems, ELTE Eötvös Loránd University(ELTE罗兰大学复杂物理系)
;
The Palynological Laboratory at the Swedish Museum of Natural History(瑞典自然历史博物馆孢粉学实验室)
;
National Centre for Public Health and Pharmacy(国家公共卫生与药品中心)
;
INRAE, UR 546 BioSP, Site Agroparc(法国国家农业、食品与环境研究院,UR 546 BioSP,阿格罗帕克园区)
;
National Korányi Institute for Pulmonology(国家科拉尼肺病研究所)
;
Health Data Science and AI Knowledge Centre, Health Services Management Training Centre, Faculty of Health and Public Administration, Semmelweis University(塞梅维什大学健康与公共管理学院卫生服务管理培训中心健康数据科学与人工智能知识中心)
;
Department of Biological Physics, ELTE Eötvös Loránd University(ELTE罗兰大学生物物理系)
专题命中
多模态评测
:multimodal(title,abstract);分类 cs.CV
AI总结
提出百万级多模态花粉显微镜数据集Pollen AI Atlas,结合专家引导的视觉-语言模型生成形态描述,实现跨区域、跨设置的高精度花粉识别与检索。
Comments31 pages, 5 main figures, supplementary information included. Submitted to Scientific Reports. v2: clarified reporting of taxonomic scope, captioning settings, backbone configuration, and evaluation details; no changes to numerical results or conclusions
机构
*
Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
;
Department of Pathology, Nanfang Hospital, Southern Medical University(南方医科大学南芳医院病理科)
;
Department of Pathology, School of Basic Medical Sciences, Southern Medical University(南方医科大学基础医学学院病理科)
;
Department of Anatomical and Cellular Pathology, Chinese University of Hong Kong(香港中文大学解剖与细胞病理学系)
;
Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理学重点实验室)
;
Jinfeng Laboratory(锦风实验室)
;
Department of Chemical and Biological Engineering, Hong Kong University of Science and Technology(香港科技大学化学与生物工程系)
;
Division of Life Science, Hong Kong University of Science and Technology(香港科技大学生命科学系)
;
State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室)
;
HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology(香港科技大学深圳-香港协同创新研究院)
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
;
Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care
RESPClinBench:呼吸专科医疗中的多模态临床决策与纵向疾病管理基准测试
Mouxiao Bian, Zhi Chen, Ruiyao Chen, Lu Lu, Hengrui Liang, Chaoyi Huang, Yiluo Lin, Jingru Ding, Yun Zhong, Yueming Su, Jie Xu
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Macau University of Science and Technology(澳门科技大学)
;
First Affiliated Hospital of Guangzhou Medical University(广州医科大学附属第一医院)
;
Guangzhou Institute of Respiratory Health(广州呼吸健康研究院)
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
LiveXiv —— 基于 arXiv 论文内容的多模态实时基准测试集
Nimrod Shabtay, Felipe Maia Polo, Sivan Doveh, Wei Lin, M. Jehanzeb Mirza, Leshem Choshen, Mikhail Yurochkin, Yuekai Sun, Assaf Arbelle, Leonid Karlinsky, Raja Giryes
机构
*
Faculty of Engineering Tel-Aviv University(特拉维夫大学工程学院)
;
IBM Research(IBM研究院)
;
Department of Statistics, University of Michigan(密歇根大学统计学系)
;
JKU Linz, Austria(林茨大学,奥地利)
;
MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
;
MIT-IBM(麻省理工学院-IBM)
Missing-by-Design: Certifiable Modality Deletion for Revocable Multimodal Sentiment Analysis
缺失-by-设计:可撤销多模态情感分析的可验证模态删除
Rong Fu, Ziming Wang, Chunlei Meng, Jiekai Wu, Kangan Qian, Hao Zhang, Simon Fong
机构
*
University of Macau(澳门大学)
;
Zhejiang University(浙江大学)
;
Fudan University(复旦大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Juntendo University(立命馆大学)
;
Tsinghua University(清华大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
Comments21 pages, 6 figures. In the previous version, Juntendo University was erroneously listed as the affiliation; we must clarify that this paper has absolutely no relation to Juntendo University. Therefore, we have replaced this affiliation in the new version
On-Device Inference versus Wireless Streaming: Energy-Efficient Multi-Modal Deep Learning for Wearable Cardiovascular Patches
面向心血管传感器贴片的端到端多模态微型CNN原型设计
Mustafa Fuad Rifet Ibrahim, Tunc Alkanat, Felix Manthey, Maurice Meijer, Alexander Schlaefer, Peer Stelldinger
机构
*
CTO System Innovation, NXP Semiconductors Germany GmbH(NXP半导体德国系统创新部)
;
Advanced Chip Engineering, NXP Semiconductors(NXP半导体先进芯片工程部)
;
Business Line Secure Connected Edge, NXP Semiconductors(NXP半导体安全连接边缘业务线)
;
Institute of Medical Technology and Intelligent Systems, Hamburg University of Technology(汉堡技术大学医学技术与智能系统研究所)
;
Department of Informatics, Hamburg University of Applied Sciences(汉堡应用科学大学信息学院)
Comments16 pages, 2 figures. Extended version of our 2024 IEEE PerCom paper, with direct on-device energy measurements, a BLE communication benchmark, architecture comparisons, and an extended evaluation. Submitted to Pervasive and Mobile Computing; Measurement-method clarifications and minor editorial corrections; results and conclusions unchanged
Ruiqi Wu, Yuang Yao, Tengfei Ma, Chenran Zhang, Na Su, Tao Zhou, Geng Chen, Wen Fan, Yi Zhou
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Department of Ophthalmology, The First Affiliated Hospital of Nanjing Medical University(南京医科大学第一附属医院眼科学系)
;
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
;
School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)
CommentsDataset can be accessed via zenodo DOI https://doi.org/10.5281/zenodo.18483292 For citation use the primary academic reference: Garbe J et al. Towards predicting sedation depth in endoscopy with large clinically annotated EEG data of continuous Propofol sedation. In: P. Andreevetal (Eds.): AIME2026, LNAI 16749, p.1-6, Springer, 2026. https://doi.org/10.1007/978-3-032-30813-9_58
PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
PixDLM:一种用于无人机推理分割的双路径多模态语言模型
Shuyan Ke, Yifan Mei, Changli Wu, Yonghan Zheng, Jiayi Ji, Liujuan Cao, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学)
Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging
学习用于多模态神经影像的稀疏潜在预测基础模型
Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian
机构
*
New York University, Center for Data Science(纽约大学数据科学中心)
;
NYU Grossman School of Medicine, Department of Radiology(纽约大学格罗斯曼医学院放射学系)
;
State University of New York at Binghamton, School of Computing(纽约州立大学宾汉姆顿分校计算机学院)
;
NYU Grossman School of Medicine, Department of Neurology(纽约大学格罗斯曼医学院神经病学系)
;
NYU Grossman School of Medicine, Department of Neurosurgery(纽约大学格罗斯曼医学院神经外科学系)
;
NYU Grossman School of Medicine, Department of Pathology(纽约大学格罗斯曼医学院病理学系)
;
School of Medicine, Department of Radiology, Stanford(斯坦福大学医学院放射学系)
;
NYU Grossman School of Medicine, Department of Neuroscience(纽约大学格罗斯曼医学院神经科学系)
;
NYU Grossman School of Medicine, Neuroscience Institute(纽约大学格罗斯曼医学院神经科学研究所)
Paired Uterine Whole-Slide Images and Pathology Reports for Multimodal Computational Pathology
用于多模态计算病理学的配对子宫全切片图像和病理报告
Han Li, Jingsong Liu, Ayako Ura, Junlin Hou, Zhengyang Xu, Azar Kazemi, Oskar Thaeter, Christian Grashei, Fabian Gülhan, Reza Nasirigerdeh, Xun Ma, Rui Yan, Hao Chen, S. Kevin Zhou, Nassir Navab, Carolin Mogler, Peter Schüffler
机构
*
Institute of Pathology, Technical University of Munich(慕尼黑工业大学病理研究所)
;
Computer Aided Medical Procedures (CAMP), Technical University of Munich(慕尼黑工业大学计算机辅助医疗程序(CAMP))
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Department of Human Pathology, Juntendo University Graduate School of Medicine(顺天堂大学医学研究生院人体病理学部)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Munich Data Science Institute (MDSI)(慕尼黑数据科学研究所)