Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
评估热门好莱坞电影的多模态叙事理解能力
David Bamman, Kent K. Chang, Allison Cooper, Juishan Hsu, Reina Kushihashi, Madison Mar, Arnav Podichetty, Rachael Samberg, Ipek Nil Sancak, Yuhan Shao
Query-Driven Multimodal Information Extraction from Long Documents
面向长文档的查询驱动多模态信息抽取
Yikai Gao, Ding Xia, Xi Yang
机构
*
School of Artificial Intelligence, Jilin University(吉林大学人工智能学院)
;
Graduate School of Information Science and Technology, The University of Tokyo(东京大学信息科学与技术研究生院)
机构
*
Microsoft Research India(微软研究院印度分部)
;
LinkedIn(领英公司)
;
Nutanix(Nutanix公司)
;
Indian Institute of Science(印度科学学院)
;
Apple(苹果公司)
;
Fujitsu Research India(富士通印度研究院)
;
Indian Institute of Technology Kharagpur(印度克勒格布尔印度理工学院)
Qixiang Yin, Huanjin Yao, Yuchen Cai, Jianghao Chen, Ziyi Wang, Min Yang, Fei Su, Zhicheng Zhao
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
ByteDance(字节跳动)
;
USTC(中国科学技术大学)
;
Beijing Key Laboratory of Network System and Network Culture(北京网络系统与网络文化重点实验室)
;
Key Laboratory of Interactive Technology and Experience System, Ministry of Culture and Tourism(文化和旅游部互动技术与体验系统重点实验室)
;
Zhongguancun Academy(中关村科学城)
机构
*
National University of Singapore(新加坡国立大学)
;
PuzzleLogic Pte Ltd(拼图逻辑私人有限公司)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳校区)
;
Peking Union Medical College Hospital(北京协和医院)
Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Training-Free Reasoning Framework
HM-Bench:多模态大语言模型在高光谱遥感中的综合基准
Xinyu Zhang, Zurong Mai, Qingmei Li, Xiaoya Fan, Zjin Liao, Haoyuan Liang, Yibin Wen, Yuhang Chen, Chan Tsz Ho, Bi Tianyuan, Ruifeng Su, Zihao Qiang, Juepeng Zheng, Jianxi Huang, Yutong Lu, Haohuan Fu
机构
*
Sun Yat-sen University(中山大学)
;
Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
;
China Agricultural University(中国农业大学)
;
Southwest Jiaotong University(西南交通大学)
;
Southwest University(西南大学)
;
National Supercomputing Center in Shenzhen(国家超级计算深圳中心)
An end-to-end-trained vision-language model for native-language prostate pathology report generation
用于生成本土语言前列腺病理报告的端到端训练视觉-语言模型
Christian Grashei, Fabian Gülhan, Maximilian Legnar, Fabian Stögbauer, Cleo-Aron Weis, Carolin Mogler, Peter Schüffler
机构
*
Technical University of Munich(慕尼黑工业大学)
;
Munich Data Science Institute(慕尼黑数据科学研究所)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
University Hospital Heidelberg(海德堡大学医院)
;
Heidelberg University(海德堡大学)
;
Interdisciplinary Center for Scientific Computing (IWR)(跨学科科学计算中心(IWR))
Comments19 pages (10 pages main text + appendix), 10 figures. Accepted at the 34th ACM International Conference on Multimedia (ACM MM 2026), Rio de Janeiro, Brazil, November 10--14, 2026
Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label
适配密集视觉-语言关系的部分标签多标签分类方法
Cheng Chen, Yifan Zhao, Jia Li
机构
*
State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室)
;
School of Computer Science and Engineering(计算机科学与工程学院)
;
Beihang University(北京航空航天大学)
Localization-Infused Vision-Language Semantic Fusion for Text-Guided Medical Image Segmentation
用于文本引导医学图像分割的定位注入视觉语言语义融合
Songyue Han, Mingye Zou, Shuchang Ye, Lei Bi, Mingyuan Meng
机构
*
Air Force Engineering University(空军工程大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
University of Sydney(悉尼大学)
;
Institute of Translational Medicine, Shanghai Jiao Tong University(上海交通大学转化医学研究院)
;
Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)
Comments12 pages, 6 figures, 7 tables. v2: revised presentation with updated author affiliations; abstract condensed; Figure 1 redrawn; an average performance column added to Table I; one reference added and three removed; author biographies removed. All datasets, methods and experimental results are unchanged from v1