Lianghuan Huang, Yihao Li, Saeed Salehi, Yingshan Chang, Ansh Soni, Konrad P. Kording
机构
*
Department of Physics and Astronomy, University of Pennsylvania, Philadelphia, PA, USA(物理与天文学系,宾夕法尼亚大学)
;
Department of Neuroscience, University of Pennsylvania, Philadelphia, PA, USA(神经科学系,宾夕法尼亚大学)
;
Department of Computer and Information Science, University of Pennsylvania(计算机与信息科学系,宾夕法尼亚大学)
;
Department of Psychology, University of Pennsylvania(心理学系,宾夕法尼亚大学)
;
Machine Learning Group, Technical University of Berlin, Berlin, Germany(机器学习组,柏林技术大学)
;
Language Technology Institute, Carnegie Mellon University, Pittsburgh, PA, USA(语言技术研究所,卡内基梅隆大学)
机构
*
Institute of Image Processing and Pattern Recognition, School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Shanghai, China(图像处理与模式识别研究所,自动化与智能感知学院,上海交通大学,上海,中国)
;
Shanghai Key Laboratory of Flexible Medical Robotics, Tongren Hospital, Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China(柔性医疗机器人上海市重点实验室,同仁医院,医疗机器人研究所,上海交通大学,上海,中国)
Attend to Anything: Foundation Model for Unified Human Attention Modeling
关注一切:统一人类注意力建模的基础模型
Wenzhuo Zhao, Ronghao Xian, Keren Fu, Qijun Zhao
机构
*
College of Computer Science, Sichuan University, Chengdu, 610065, China(四川大学计算机学院)
;
National Key Laboratory of Fundamental Science on Synthetic Vision, Sichuan University, Chengdu, 610065, China(合成视觉基础科学国家重点实验室)
AI总结
提出 Attend to Anything Model (AAM),一种多模态基础模型,通过层次化语言提示和双曲空间嵌入统一图像、视频和视听任务中的注意力建模,并在16个基准上平均提升6%,视频推理加速约4倍。