DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri, Thomas Looi, Dale Podolsky, Mehrdad Zadeh, Javad Dargahi
机构
*
Gina Cody School of Engineering and Computer Science, Concordia University(甘娜·柯迪工程与计算机科学学院,康科迪亚大学)
;
The Wilfred and Joyce Posluns Centre for Image Guided Innovation & Therapeutic Intervention (PCIGITI) at the Hospital for Sick Children (SickKids)(威廉与乔伊斯·波斯卢斯影像引导创新与治疗干预中心(PCIGITI)(SickKids医院))
;
Electrical and Computer Engineering Department, Kettering University(电气与计算机工程系,凯特林大学)
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan Pravin More, Sanoojan Baliah, Hasindri Watawana, Yuhao Li, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber, Hisham Cholakkal, Ivan Laptev, Shin'ichi Satoh, Michael Felsberg, Mubarak Shah, Salman Khan, Fahad Shahbaz Khan
机构
*
Mohamed bin Zayed University of Artificial Intelligence(莫德赫·本·扎耶德人工智能大学)
;
University of Central Florida(中央佛罗里达大学)
;
Islamic University of Technology(伊斯兰技术大学)
;
Air University(空军大学)
;
ETH Zurich(苏黎世联邦理工学院)
;
Technische Universität München(慕尼黑技术大学)
;
National Institute of Informatics(国家信息研究所)
;
Australian National University(澳大利亚国立大学)
;
Linköping University(利尔贝里大学)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
Zhen Chen, Xingjian Luo, Kun Yuan, Jinlin Wu, Danny T. M. Chan, Nassir Navab, Hongbin Liu, Zhen Lei, Jiebo Luo
机构
*
Hong Kong Institute of Science & Innovation(香港科学与工业创新研究院)
;
CAMP, Technische Universität München(CAMP,慕尼黑技术大学)
;
Department of Surgery, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院外科部)
机构
*
National University of Singapore(新加坡国立大学)
;
Peking University(北京大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Estates Pte Ltd(6Estates私人有限公司)
Bing Wang, Ximing Li, Mengzhe Ye, Changchun Li, Bo Fu, Jianfeng Qu, Lin Yuanbo Wu
机构
*
College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院)
;
College of Software, Jilin University(吉林大学软件学院)
;
School of Computer and Artificial Intelligence, Liaoning Normal University(辽宁师范大学计算机与人工智能学院)
;
School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)
;
Department of Computer Science, Swansea University(斯旺西大学计算机科学系)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
Kun Yuan, Vinkle Srivastav, Tong Yu, Joel L. Lavanchy, Jacques Marescaux, Pietro Mascagni, Nassir Navab, Nicolas Padoy
机构
*
University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France(斯特拉斯堡大学,法国国家科学研究中心,法国国家卫生研究院,ICube,UMR7357,法国斯特拉斯堡)
;
University Digestive Health Care Center – Clarunis, 4002 Basel, Switzerland(消化健康研究中心–Clarunis,瑞士巴塞尔)
机构
*
College of Computer and Information Engineering, Tianjin Normal University(天津师范大学计算机与信息工程学院)
;
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳))
;
College of Intelligence and Computing, Tianjin University(天津大学智能科学与计算学院)
;
The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江省区块链与数据安全国家重点实验室)
;
Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, Hangzhou(杭州高新技术区(滨江)区块链与数据安全研究院)