DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri, Thomas Looi, Dale Podolsky, Mehrdad Zadeh, Javad Dargahi
机构
*
Gina Cody School of Engineering and Computer Science, Concordia University(甘娜·柯迪工程与计算机科学学院,康科迪亚大学)
;
The Wilfred and Joyce Posluns Centre for Image Guided Innovation & Therapeutic Intervention (PCIGITI) at the Hospital for Sick Children (SickKids)(威廉与乔伊斯·波斯卢斯影像引导创新与治疗干预中心(PCIGITI)(SickKids医院))
;
Electrical and Computer Engineering Department, Kettering University(电气与计算机工程系,凯特林大学)
CommentsThe 1st place report of 7th LSVOS challenge RVOS track in ICCV 2025. The code is released in Sa2VA repository: https://github.com/bytedance/Sa2VA
RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
Kunyu Peng, Di Wen, Jia Fu, Jiamin Wu, Kailun Yang, Junwei Zheng, Ruiping Liu, Yufan Chen, Yuqian Fu, Danda Pani Paudel, Luc Van Gool, Rainer Stiefelhagen
机构
*
Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology(人机化研究所,卡尔斯鲁厄技术大学)
;
RISE Research Institutes of Sweden(瑞典RISE研究机构)
;
KTH Royal Institute of Technology(皇家理工学院)
;
School of Artificial Intelligence and Robotics(人工智能与机器人学院)
;
National Engineering Research Center of Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心)
;
Chinese University of Hong Kong(香港中文大学)
;
Shanghai AI Lab(上海人工智能实验室)
机构
*
Tencent(腾讯公司)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Zhejiang University(浙江大学)
;
Singapore University of Technology and Design(新加坡科技设计大学)
Towards Interpretable and Trustworthy Time Series Reasoning: A BlueSky Vision
Kanghui Ning, Zijie Pan, Yushan Jiang, Anderson Schneider, Yuriy Nevmyvaka, Dongjin Song
机构
*
School of Computing University of Connecticut Storrs, CT(计算学院 美国康涅狄格大学 斯托尔斯分校)
;
Department of Machine Learning Research Morgan Stanley New York, NY(机器学习研究部 花旗集团 新 York)