COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
COMPASS:在统一多模态模型中锚定构图意图引导
Ziqi Zhou, Weize Quan, Mining Tan, Zhihan Chen, Dandan Zheng, Jingdong Chen, Jun Zhou, Weiming Dong, Dong-Ming Yan
机构
*
University of Edinburgh(爱丁堡大学)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS)(多模态人工智能系统国家重点实验室)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Ant Group(蚂蚁集团)
Weihao Tan, Changjiu Jiang, Yu Duan, Mingcong Lei, Jiageng Li, Yitian Hong, Xinrun Wang, Bo An
机构
*
Nanyang Technological University(南洋理工大学)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
East China University of Science and Technology(华东理工大学)
;
Singapore Management University(新加坡管理大学)
Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training
跨视角多模态视觉评估框架用于中医康复训练
Francis Xiatian Zhang, Hao Yao, Shengxuan Chen, Hong Zhu, Hongxiao Jia, Sisi Zheng, Hubert P. H. Shum
机构
*
Department of Computer Science, Durham University(杜伦大学计算机科学系)
;
Institute for Regeneration and Repair, The University of Edinburgh(爱丁堡大学再生与修复研究所)
;
Ningbo Hospital of Traditional Chinese Medicine(宁波市中医院)
;
Department of Rehabilitation Medicine, The Gulou Hospital of Traditional Chinese Medicine(南京市鼓楼区中医院康复医学科)
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
QiYuanLab(启元实验室)
;
Tsinghua University(清华大学)
;
University of Electronic Science and Technology of China(电子科技大学)
机构
*
University of International Relations(国际关系学院)
;
Kedge Business School(凯致商学院)
;
Peking University(北京大学)
;
Tsinghua University(清华大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Capital Normal University(首都师范大学)
机构
*
organization= Department of Geography, Texas A\&M University , city= College Station , country= USA
;
organization= Department of Landscape Architecture \& Urban Planning, Texas A\&M University , city= College Station , country= USA
;
organization= Department of Industrial
;
Systems Engineering, University of Florida , city= Gainesville , country= USA
;
organization= Spatial Sciences Institute, University of Southern California , city= Los Angeles , country= USA
;
organization= Department of Geography
;
Sustainability, University of Tennessee , city= Knoxville , country= USA
HEad and neCK TumOR (HECKTOR) 2025: Benchmark of Segmentation, Diagnosis, and Prognosis in Multimodal PET/CT
头颈肿瘤 (HECKTOR) 2025 挑战赛:多模态 PET/CT 中的分割、诊断与预后基准
Numan Saeed, Salma Hassan, Shahad Hardan, Lishan Cai, Xinglong Liang, Moona Mazher, Abdul Qayyum, Yansong Bu, Mengye Lyu, Yue Lin, Mingyuan Meng, Chuanyi Huang, Lisheng Wang, Dalal Chamseddine, Shamimeh Ahrari, Beining Wu, Yifei Chen, Fuyou Mao, Hao Zhang, Baixiang Zhao, Surajit Ray, Muzi Guo, Lei Xiang, Jakob Dexl, Michael Ingrisch, Adrien Depeursinge, Arman Rahmim, Mathieu Hatt, Vincent Andrearczyk, Mohammad Yaqub
机构
*
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
;
Amsterdam UMC(阿姆斯特丹大学医学中心)
;
The Netherlands Cancer Institute(荷兰癌症研究所)
;
Radboud University Medical Centre(拉德堡德大学医学中心)
;
University College London(伦敦大学学院)
;
Imperial College London(帝国理工学院)
;
Shenzhen Technology University(深圳技术大学)
;
Shenzhen University(深圳大学)
;
Newland Digital Technology(新大陆数字技术)
;
The University of Sydney(悉尼大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University Hospital, Nantes(南特大学医院)
;
Nantes Université, Centrale Nantes, CNRS, LS2N(南特大学、南特中央理工学院、法国国家科学研究中心、LS2N实验室)
;
Hangzhou Dianzi University(杭州电子科技大学)
;
Tsinghua University(清华大学)
;
Central South University(中南大学)
;
University of Glasgow(格拉斯哥大学)
;
China Mobile System Integration Co., Ltd.(中移系统集成有限公司)
;
Subtle Medical Inc.(Subtle Medical公司)
;
University Hospital, LMU Munich(慕尼黑大学医院)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
BC Cancer Research Institute(不列颠哥伦比亚癌症研究所)
;
HES-SO Valais-Wallis University of Applied Sciences and Arts(HES-SO瓦莱州应用科学与艺术大学)
;
Lausanne University Hospital (CHUV)(洛桑大学医院)
;
LaTIM, INSERM, UMR 1101, Univ Brest(LaTIM实验室、法国国家健康与医学研究院、UMR 1101、布雷斯特大学)
Comments17 pages, 4 figures, 4 tables. Overview paper for the HECKTOR 2025 challenge, held as a satellite event at MICCAI 2025. Challenge website: https://hecktor.grand-challenge.org/
机构
*
Jülich Supercomputing Centre (JSC), Forschungszentrum Jülich(julich超级计算中心(JSC),julich研究所)
;
School of Engineering and Natural Sciences (SENS), University of Iceland(工程与自然科学学院(SENS),冰岛大学)
;
Global Land Monitoring Group, GFZ Helmholtz Centre for Geosciences(全球土地监测组,geofz赫尔姆霍兹研究中心)
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning
MathVis-Fine:通过渐进式依赖引导训练将视觉监督与必要性对齐的多模态数学推理
Wanshi Xu, Haokun Zhao, Haidong Yuan, Songjun Cao, Long Ma
机构
*
School of ECE, Peking University(北京大学电子与计算机工程学院)
;
College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与技术学院)
;
School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
;
Tencent Youtu Lab(腾讯优图实验室)
Comments8 pages, 6 figures. To appear in Proceedings of the 8th International Workshop on IoT Applications and Industry 5.0 (IoTI5 2026), co-located with IEEE DCOSS-IoT 2026, Reykjavik, Iceland, June 2026
Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis
探测、融合与可信度:基础模型表示在多模态癌症分析中的系统评估
Jingyu Hu, Giuseppe Tripodi, Reed Naidoo, Sarah F. McGough, Tapabrata Chakraborti
机构
*
The Alan Turing Institute(艾伦·图灵研究所)
;
University of Bristol(布里斯托大学)
;
University of Manchester(曼彻斯特大学)
;
The Institute of Cancer Research(癌症研究所)
;
Genentech(基因泰克)
UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics
UrbanWell: 面向时空城市福祉分析的多模态大语言模型基准测试
Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui
机构
*
University of Helsinki(赫尔辛基大学)
;
Zhongguancun Academy(中关村学院)
;
University of Oxford(牛津大学)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens
OmniMouse: 基于1500亿神经令牌的多模态多任务脑模型的可扩展性
Konstantin F. Willeke, Polina Turishcheva, Alex Gilbert, Goirik Chakrabarty, Hasan A. Bedel, Paul G. Fahey, Yongrong Qiu, Marissa A. Weis, Michaela Vystrčilová, Taliah Muhammad, Lydia Ntanavara, Rachel E. Froebe, Kayla Ponder, Zheng Huan Tan, Emin Orhan, Erick Cobos, Sophia Sanborn, Katrin Franke, Fabian H. Sinz, Alexander S. Ecker, Andreas S. Tolias
机构
*
Department of Ophthalmology, Byers Eye Institute, Stanford University(斯坦福大学眼科学系、比尔斯眼科研究所)
;
Stanford Bio-X, Stanford University(斯坦福大学生物交叉学科)
;
Wu Tsai Neurosciences Institute, Stanford University(斯坦福大学吴泰教授神经科学研究所)
;
Institute of Computer Science and Campus Institute Data Science, University Göttingen(哥廷根大学计算机科学研究所和校园数据科学研究所)