SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
机构 * Columbia University(哥伦比亚大学) ; University of Washington(华盛顿大学)
高校专区
机构 * Columbia University(哥伦比亚大学) ; University of Washington(华盛顿大学)
机构 * Department of Electrical Engineering, Columbia University(哥伦比亚大学电气工程系)
Comments EMNLP 2025 Main Conference (Oral)
机构 * Carnegie Mellon University(卡内基梅隆大学) ; University of Notre Dame(圣母大学) ; Columbia University(哥伦比亚大学)
机构 * Department of Biomedical Informatics, Stony Brook University(生物医学信息学系,石溪大学) ; Department of Radiology, Columbia University Irving Medical Center(放射学系,哥伦比亚大学伊万杰琳医疗中心) ; Department of Neuro-Oncology, Columbia University Irving Medical Center(神经肿瘤学系,哥伦比亚大学伊万杰琳医疗中心)
机构 * Columbia University(哥伦比亚大学) ; Brown University(布朗大学) ; Purdue University(普渡大学) ; Google(谷歌)
Comments Extended version of paper to be published in the proceedings of ACM CCS 2025
机构 * Department of Computer Science, Columbia University(哥伦比亚大学计算机科学系) ; Lincoln Laboratory, Massachusetts Institute of Technology(麻省理工学院林肯实验室) ; Paul G. Allen School of Computer Science & Engineering, University of Washington(华盛顿大学保罗·G·艾伦计算机科学与工程学院) ; Barnard College, Columbia University(哥伦比亚大学巴纳德学院)
Comments 8 pages, 7 figures, accepted to IROS 2025