Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization
Sooyoung Park, Arda Senocak, Joon Son Chung
机构
*
School of Electrical Engineering, KAIST, South Korea(韩国延世大学电气工程学院)
;
ETRI, South Korea(韩国电子技术研究院)
专题命中
音频语音多模态
:multimodal(abstract);audio-visual(abstract);multimodal foundation model(abstract);分类 cs.CV、eess.AS
CommentsJournal Extension of WACV 2024 paper (arXiv:2311.04066). Code is available at ACL-SSL" target="_blank" rel="noopener">https://github.com/swimmiing/ACL-SSL
Legible and Intuitive Multi-modal Robot State and Intent Communication Validated in Online and Real-world Studies
可读且直观的多模态机器人状态与意图通信:在线和真实世界研究验证
Tim Schreiter, Jens V. Rüppel, Andrey Rudenko, Martin Magnusson, Achim J. Lilienthal
机构
*
Chair of Perception for Intelligent Systems, Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM)(慕尼黑工业大学慕尼黑机器人与机器智能研究所智能系统感知教席)
;
Centre for Applied Autonomous Sensor Systems (AASS), Örebro University(厄勒布鲁大学应用自主传感器系统中心)
;
Robotics Institute Germany (RIG)(德国机器人研究所)
Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models
使用基于语音的多模态大语言模型实现可泛化的认知障碍检测
Yingchao Huang, Xin Wang, Yuhan Su, Shanshan Yao
机构
*
Faculty of Digital Innovation, Arts \& Sciences, Saskatchewan Polytechnic, Regina SK S4S 5X1, Canada
;
School of Basic Medical Sciences, Hebei University, Baoding 071000, China
;
Department of Civil \& Environmental Engineering
;
School of Mining \& Petroleum Engineering, University of Alberta, Edmonton AB T6G 2H5, Canada
CommentsSubmitted to 2026 Joint 14th International Conference on Soft Computing and Intelligent Systems and 27th International Symposium on Advanced Intelligent Systems (SCIS&ISIS 2026)
Audio-Visual Speech Enhancement: Architectural Design and Deployment Strategies
音频-视觉语音增强:架构设计与部署策略
Anis Hamadouche, Haifeng Luo, Mathini Sellathurai, Amir Hussain, Tharm Ratnarajah
机构
*
School of Engineering & Physical Sciences, Heriot-Watt University(赫瑞斯泰学院,赫瑞斯泰大学)
;
College of Engineering Department of Electrical and Computer Engineering, San Diego State University(工程学院电子与计算机工程系,圣地亚哥州立大学)
;
SDAIA-KFUPM Joint Research Centre for Artificial Intelligence, King Fahd University of Petroleum and Minerals(SDAIA-KFUPM人工智能联合研究中心,国王法赫德石油与矿物大学)
Scaling to Multimodal and Multichannel Heart Sound Classification with Synthetic and Augmented Biosignals
利用合成与增强生物信号实现多模态和多通道心音分类的规模化
Milan Marocchi, Matthew Fynn, Kayapanda Mandana, Yue Rong
机构
*
School of Electrical Engineering, Computing, and Mathematical Sciences (EECMS), Faculty of Science and Engineering, Curtin University, Bentley, WA 6102, Australia(电气工程、计算与数学科学学院(EECMS),科学与工程学院, Curtin 大学,Bentley,WA 6102,澳大利亚)