Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
用于以音频为中心任务的音频-语言模型:系统综述
Yi Su, Jisheng Bai, Qisheng Xu, Kele Xu, Yong Dou
机构
*
College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学)
;
School of Communications and Information Engineering, Xi’an University of Posts and Telecommunications(通信与信息工程学院,西安邮电大学)
Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
视频中矛盾/犹豫识别用于个性化数字健康干预
Manuela González-González, Soufiane Belharbi, Muhammad Osama Zeeshan, Masoumeh Sharafi, Muhammad Haseeb Aslam, Lorenzo Sia, Nicolas Richet, Marco Pedersoli, Alessandro Lameiras Koerich, Simon L Bacon, Eric Granger
机构
*
LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(ETS蒙特利尔大学系统工程系LIVIA实验室)
;
LIVIA, Dept. of Software and IT Engineering, ETS Montreal, Canada(ETS蒙特利尔大学软件与信息工程系LIVIA实验室)
;
Dept. of Health, Kinesiology, & Applied Physiology, Concordia University, Montreal, Canada(康科迪亚大学健康、运动科学与应用生理学系)
;
Montreal Behavioural Medicine Centre, CIUSSS Nord-de-l’Ile-de-Montréal, Canada(蒙特利尔行为医学中心,蒙特利尔北岛卫生与社会服务局)
机构
*
South China University of Technology(华南理工大学)
;
The University of Hong Kong(香港大学)
;
Snap Inc(Snap公司)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
GuideMe:流视频中的多域任务指导与干预
Fang Liu, Jinpeng Chen, Ke Xu, Yuhao Liu, Huankang Guan, Xudong Lu, Bo Yang, Gerhard Hancke, Rui Liu, Rynson W. H. Lau
机构
*
City University of Hong Kong(香港城市大学)
;
Huawei Research(华为研究院)
;
University of Science and Technology of China(中国科学技术大学)
;
Chinese University of Hong Kong(香港中文大学)
;
City University of Hong Kong (Dongguan)(香港城市大学(东莞))
CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors
CorridorVLA:通过稀疏锚点实现生成动作头的显式空间约束
Dachong Li, ZhuangZhuang Chen, Jin Zhang, Jianqiang Li
机构
*
College of Computer Science and Software Engineering(计算机科学与软件工程学院)
;
National Engineering Laboratory for Big Data System Computing Technology(大数据系统计算技术国家工程实验室)
Comments12 pages, 2 figures. v2: adds equivalence and Bayes-factor bounds, split-half reliability ceiling, a supervised probe with proper position control, a per-input-stream dissociation, and ISC discussion. Code, video-ID manifest, and per-video results: https://github.com/mercurialsolo/tribe-replay-heatmaps
CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation
CoLA-Flow Policy: 通过连续潜在动作流匹配实现机器人操作的时序一致模仿学习
Wu Songwei, Jiang Zhiduo, Sun Wandong, Xie Guanghu, Zhao Rui, Liu Hong, Liu Yang
机构
*
State Key Laboratory of Robotics and System, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)
;
The University of Sydney(悉尼大学)
;
Honor Device Co., Ltd.(荣耀设备有限公司)