Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge
超越语言和模态限制学习说话者身份:来自POLY-SIM 2026挑战赛的见解
Marta Moscati, Muhammad Saad Saeed, Marina Zanoni, Mubashir Noman, Rohan Kumar Das, Monorama Swain, Yassin Terraf, Yufang Hou, Elisabeth Andre, Khalid Mahmood Malik, Markus Schedl, Shah Nawaz
机构
*
Institute of Computational Perception, Johannes Kepler University(约翰内斯·开普勒大学计算感知研究所)
;
University of Michigan-Flint(密歇根大学弗林特分校)
;
BDO DIGITAL Gmbh(BDO数字有限公司)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Fortemedia(富特媒体公司)
;
College of Computing, Mohammed VI Polytechnic University(穆罕默德六世理工大学计算学院)
;
IT:U Interdisciplinary Transformation University(IT:U跨学科转型大学)
;
University of Augsburg(奥格斯堡大学)
机构
*
School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络空间安全学院)
;
College of Computer Science, Chongqing University(重庆大学计算机科学学院)
;
School of Software and engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)
Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference
眼见无需耗能,言语却要代价:揭示边缘视觉语言模型推理中的真正能量瓶颈
Junfei Zhan, Haoxun Shen, Mingang Guo, Zixuan Huang, Tengjiao He
机构
*
University of Pennsylvania(宾夕法尼亚大学)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
Jinan University(暨南大学)