机构
*
School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院)
;
College of Computer Science, Nankai University(南开大学计算机学院)
;
School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
;
Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室)
;
National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)
专题命中
图文多模态
:multimodal(title,abstract);分类 cs.CV
CommentsAccepted by IEEE Transactions on Instrumentation and Measurement (TIM)
Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models
Shuang Liang, Zhihao Xu, Jialing Tao, Hui Xue, Xiting Wang
专题命中
图文多模态
:multi-modal(abstract);分类 cs.CV、cs.AI
CommentsWithdrawn due to an accidental duplicate submission. This paper (arXiv:2510.15430) was unintentionally submitted as a new entry instead of a new version of our previous work (arXiv:2508.09201)
机构
*
Department of Computer Vision(计算机视觉系)
;
Mohamed bin Zayed University of Artificial Intelligence(马尔代夫比兹人工智能大学)
;
Department of Machine Learning(机器学习系)
;
Corniche Hospital, Abu Dhabi Health Services Company (SEHA)(阿布扎赫尔医院,阿布扎赫健康服务公司(SEHA))
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
Che Liu, Yingji Zhang, Dong Zhang, Weijie Zhang, Chenggong Gong, Yu Lu, Shilin Zhou, Ziliang Gan, Ziao Wang, Haipang Wu, Ji Liu, André Freitas, Qifan Wang, Zenglin Xu, Rongjuncheng Zhang, Yong Dai
机构
*
Imperial College London(伦敦帝国学院)
;
University of Manchester(曼彻斯特大学)
;
HiThink Research(HiThink研究院)
;
Soochow University(苏州大学)
;
Hong Kong Baptist University(香港 Baptist大学)
;
Idiap Research Institute(Idiap研究 institute)
;
Meta AI
;
Fudan University(复旦大学)
DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri, Thomas Looi, Dale Podolsky, Mehrdad Zadeh, Javad Dargahi
机构
*
Gina Cody School of Engineering and Computer Science, Concordia University(甘娜·柯迪工程与计算机科学学院,康科迪亚大学)
;
The Wilfred and Joyce Posluns Centre for Image Guided Innovation & Therapeutic Intervention (PCIGITI) at the Hospital for Sick Children (SickKids)(威廉与乔伊斯·波斯卢斯影像引导创新与治疗干预中心(PCIGITI)(SickKids医院))
;
Electrical and Computer Engineering Department, Kettering University(电气与计算机工程系,凯特林大学)
CommentsThe 1st place report of 7th LSVOS challenge RVOS track in ICCV 2025. The code is released in Sa2VA repository: https://github.com/bytedance/Sa2VA
RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
Kunyu Peng, Di Wen, Jia Fu, Jiamin Wu, Kailun Yang, Junwei Zheng, Ruiping Liu, Yufan Chen, Yuqian Fu, Danda Pani Paudel, Luc Van Gool, Rainer Stiefelhagen
机构
*
Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology(人机化研究所,卡尔斯鲁厄技术大学)
;
RISE Research Institutes of Sweden(瑞典RISE研究机构)
;
KTH Royal Institute of Technology(皇家理工学院)
;
School of Artificial Intelligence and Robotics(人工智能与机器人学院)
;
National Engineering Research Center of Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心)
;
Chinese University of Hong Kong(香港中文大学)
;
Shanghai AI Lab(上海人工智能实验室)
机构
*
Tencent(腾讯公司)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Zhejiang University(浙江大学)
;
Singapore University of Technology and Design(新加坡科技设计大学)
Towards Interpretable and Trustworthy Time Series Reasoning: A BlueSky Vision
Kanghui Ning, Zijie Pan, Yushan Jiang, Anderson Schneider, Yuriy Nevmyvaka, Dongjin Song
机构
*
School of Computing University of Connecticut Storrs, CT(计算学院 美国康涅狄格大学 斯托尔斯分校)
;
Department of Machine Learning Research Morgan Stanley New York, NY(机器学习研究部 花旗集团 新 York)
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
Yuyang Hong, Jiaqi Gu, Qi Yang, Lubin Fan, Yue Wu, Ying Wang, Kun Ding, Shiming Xiang, Jieping Ye
机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS)
;
Alibaba Cloud Computing(阿里巴巴云计算)
BenCao: An Instruction-Tuned Large Language Model for Traditional Chinese Medicine
Jiacheng Xie, Yang Yu, Yibo Chen, Hanyao Zhang, Lening Zhao, Jiaxuan He, Lei Jiang, Xiaoting Tang, Guanghui An, Dong Xu
机构
*
Community Health Service Center Shanghai Pudong New Area(上海浦东新区社区卫生服务中心)
;
School of Acupuncture-Moxibustion and Tuina, Shanghai University of Traditional Chinese Medicine(上海中医药大学针灸推拿学院)
Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
Min Cao, Xinyu Zhou, Ding Jiang, Bo Du, Mang Ye, Min Zhang
机构
*
School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Key Laboratory of New Generation Artificial Intelligence Technology & Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新 generation 人工智能技术及交叉应用重点实验室(东南大学),教育部,中国)