CommentsAccepted by CVPR 2024 workshop. The 1st winning model in CVPR 2024 UG2+ challenge. The code and configuration of our method are available at https://github.com/dtc111111/Multi-Modal-UAV
CommentsPresented at ACAIN 2025 (Advanced Course & Symposium on Artificial Intelligence & Neuroscience), recipient of the Camillo Golgi Best Paper Award
MeMo: Attentional Momentum for Real-Time Audio-Visual Target Speaker Extraction Under Impaired Visual Conditions
MeMo: 视觉受损条件下的实时视听目标说话人提取的注意力动量
Junjie Li, Wenxuan Wu, Shuai Wang, Zexu Pan, Kong Aik Lee, Helen Meng, Haizhou Li
机构
*
Department of Electrical and Electronic Engineering, Faculty of Engineering, The Hong Kong Polytechnic University(电子工程系,工程学院,香港理工大学)
;
Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong(系统工程与工程管理系,香港中文大学)
;
School of Artificial Intelligence (SAI), The Chinese University of Hong Kong, Shenzhen(人工智能学院(SAI),香港中文大学深圳校区)
;
School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学)
;
Tongyi Lab, Alibaba Group, Singapore(通义实验室,阿里巴巴集团,新加坡)
`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs
从提示到扰动:针对音频大语言模型的自适应语音越狱框架
Linghan Huang, Bo Li, Huaming Chen, Kim-Kwang Raymond Choo
机构
*
School of Electrical and Computer Engineering, The University of Sydney(悉尼大学电气与计算机工程学院)
;
University of Chicago(芝加哥大学)
;
University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)