AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance
AMAVA:一种自适应的动觉感知视频到音频框架,用于视障协助
Benjamin Klein, Kazi Ruslan Rahman, Sanchita Ghose
机构
*
Department of Computer Science, San Francisco State University(计算机科学系,旧金山州立大学)
;
Department of Mathematics, San Francisco State University(数学系,旧金山州立大学)
;
Department of Computer Engineering, San Francisco State University(计算机工程系,旧金山州立大学)
Comments8 pages, 7 figures. Published in the Proceedings of the 15th International Conference on Pattern Recognition Applications and Methods (ICPRAM 2026), pages 282--289
Journal refIn Proceedings of the 15th International Conference on Pattern Recognition Applications and Methods (ICPRAM 2026), pages 282--289, 2026
机构
*
School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
;
National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources
BhashaSutra:印度NLP数据集、语料库和资源的任务导向统一调查
Raghvendra Kumar, Devankar Raj, Sriparna Saha
机构
*
Department of Computer Science and Engineering, Indian Institute of Technology Patna, India(印度理工学院帕纳加邦计算机科学与工程系)
;
Indian Institute of Technology Patna, India(印度理工学院帕纳加邦)
FreezeEmpath: Efficient Training for Empathetic Spoken Chatbots with Frozen LLMs
FreezeEmpath: 基于冻结大语言模型的高效共情语音聊天机器人训练
Yun Hong, Yan Zhou, Yang Feng
机构
*
Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences(智能信息处理重点实验室,计算技术研究所,中国科学院)
;
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院)
;
University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国)
Acoustic and Facial Markers of Perceived Conversational Success in Spontaneous Speech
语音与面部特征在自发言语中感知对话成功的标志
Thanushi Withanage, Elizabeth Redcay, Carol Espy-Wilson
机构
*
Department of Electrical and Computer Engineering, University of Maryland College Park, MD, USA(电气与计算机工程系,马里兰大学学院市分校)
;
Department of Psychology, University of Maryland College Park, MD, USA(心理学系,马里兰大学学院市分校)
CommentsCamera-ready version. 14 pages, 5 figures in total: 8 pages main text with 4 figures, 3 pages references, and 3 pages appendix with 1 figure. Accepted at the 10th ABAW Workshop, CVPR 2026
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
MCAT: 通过MLLMs扩展多对多语音到文本翻译至70种语言
Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Keqi Deng, Xie Chen, Yang Xiang, Ming Liu, Bing Qin, YaoWei Wang
机构
*
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
Pengcheng Laboratory(鹏城实验室)
;
Harbin Institute of Technology, Harbin(哈尔滨工业大学(哈尔滨))
;
University of Cambridge(剑桥大学)
;
Shanghai Jiao Tong University(上海交通大学)
机构
*
College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与技术学院)
;
College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)
;
Jiiov Technology(极奥科技)
Comments11 pages, 9 figures. This is the author's version of the article that appeared at the IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR) 2026
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
Omni-MMSI:迈向基于身份的社会互动理解
Xinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng, Cihang Xie, Yuyin Zhou, James Matthew Rehg, Yapeng Tian
机构
*
University of Texas at Dallas(德克萨斯大学达拉斯分校)
;
Georgia Institute of Technology(佐治亚理工学院)
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)