机构
*
Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究所,东部技术研究所)
;
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Munich Center for Machine Learning, LMU Munich(慕尼黑大学机器学习中心,慕尼黑大学)
专题命中
VLM训练与架构
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
MediX-R1: Open Ended Medical Reinforcement Learning
MediX-R1:开放端医疗强化学习
Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Omair Mohamed, Mohamed Zidan, Fahad Khan, Salman Khan, Rao Anwer, Hisham Cholakkal
机构
*
Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(迈赫迈德·本·扎耶德人工智能大学)
;
Jubilee Mission Medical College(jubilee mission 医学院)
;
Research Institute(研究院)
;
JJM Medical College(JJM 医学院)
专题命中
VLM训练与架构
:VLM(abstract);multimodal large language model(abstract);分类 cs.CV
Brain3D: Brain Report Automation via Inflated Vision Transformers in 3D
Brain3D: 通过膨胀视觉变换器实现脑部报告自动化
Mariano Barone, Francesco Di Serio, Giuseppe Riccio, Antonio Romano, Marco Postiglione, Antonino Ferraro, Vincenzo Moscato
机构
*
University of Naples Federico II, Department of Electrical Engineering and Information Technology(那不勒斯费德里科二世大学电子工程与信息技术系)
;
Northwestern University, Dept. of Computer Science, McCormick School of Engineering and Applied Science(西北大学计算机科学系,工程与应用科学学院)
;
Pegaso University, Department of Information Science and Technology(佩加索大学信息科学与技术系)
机构
*
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
;
Peng Cheng Laboratory(鹏城实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Peking University(北京大学)
Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter
Mod-Adapter:通过调制适配器实现无需微调的多概念个性化
Weizhi Zhong, Huan Yang, Zheng Liu, Huiguo He, Zijian He, Xuesong Niu, Di Zhang, Guanbin Li
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
Kolors Team, Kuaishou Technology(快手科技Kolors团队)
;
Shenzhen Loop Area Institute, Shenzhen, China(深圳循环区研究所)
;
Guangdong Key Laboratory of Big Data Analysis and Processing, Guangzhou, China(广东大数据分析与处理重点实验室)
;
South China University of Technology, Guangzhou, China(华南理工大学)
CommentsThis work has been accepted to the International Conference on Learning Representations (ICLR) 2026. Project Page: https://rfdetr.roboflow.com/
EventFlash: Towards Efficient MLLMs for Event-Based Vision
EventFlash: 向基于事件的视觉高效MLLMs迈进
Shaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang, Wen Jiang, Ming Li, Xiangyang Ji
机构
*
Xidian University(西安电子科技大学)
;
Tsinghua University(清华大学)
;
Beijing Institute of Technology(北京理工大学)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy(SZ)(广东人工智能与数字经济实验室(深圳))
专题命中
VLM训练与架构
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing
MMR-Bench: 多模态大语言模型路由的综合基准
Haoxuan Ma, Guannan Lai, Han-Jia Ye
机构
*
School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学)
;
National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家实验室,南京大学)
专题命中
VLM训练与架构
:multimodal large language model(abstract);MLLM(abstract);分类 cs.AI