Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
机构 * University of California, San Diego(加州大学圣迭戈分校) ; Adobe Research(Adobe研究)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Preprint
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * University of California, San Diego(加州大学圣迭戈分校) ; Adobe Research(Adobe研究)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Preprint
机构 * AIDAS Laboratory(AIDAS实验室) ; IPAI ; ECE(电子工程系) ; Seoul National University(首尔国立大学)
专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI
Comments Accepted in NeurIPS 2025
机构 * Peter Munk Cardiac Centre, University Health Network (UHN)(彼得·默克心脏中心,大学健康网络) ; Department of Medical Biophysics, UofT(医学生物物理学系) ; Ted Rogers Centre for Heart Research, UHN(泰德·罗杰斯心脏病研究中心,大学健康网络) ; Department of Computer Science, University of Toronto (UofT)(计算机科学系,多伦多大学) ; Toronto General Hospital Research Institute, UHN(多伦多总医院研究 institute) ; Department of Medical Imaging, UofT(医学影像学系) ; Vector Institute, Toronto(向量研究所)
专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV
Comments MICCAI 2025
Journal ref Medical Image Computing and Computer Assisted Intervention - MICCAI 2025. MICCAI 2025. Lecture Notes in Computer Science, vol 15964. Springer, Cham
专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI
Comments First Peer Reviewed Review Paper for Object Detection with Vision-Language Models (VLMs)
Journal ref Information Fusion, 2025
机构 * School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院) ; College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院) ; School of Computer Science and Technology and Ministry of Education Key Lab For Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院和教育部智能网络与网络安全重点实验室) ; School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学数学与统计学院和教育部智能网络与网络安全重点实验室) ; Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China(琶洲实验室(黄埔),广州,广东,中国)
专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV
机构 * University of Manchester(曼彻斯特大学) ; South China University of Technology(华南理工大学)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments Accept by EMNLP2025
机构 * Zhuoning Xu 1(Xu Zhuoning 1) ; Xinyan Liu 1(Liu Xinyan 1)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract)
机构 * VUNO Inc.(VUNO公司) ; KAIST(韩国科学技术院)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI
Comments 38 pages, 17 figures, preprint
机构 * Birla Institute of Technology and Science Pilani, K.K Birla Goa Campus(比拉理工学院和科学学院,K.K比拉果阿校区)
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV
Journal ref Proc. European Conference on Mobile Robots (ECMR), 2025, pp. 1-6
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title,abstract);分类 cs.CV
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV
机构 * Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center, School of Computer Science and Technology, Xinjiang University(新疆多模态智能处理与信息安全工程技术创新中心,计算机科学与技术学院,新疆大学) ; Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) ; School of Electrical Engineering and Automation, Tianjin University of Technology(电气工程与自动化学院,天津工业大学)
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.AI、cs.MM
Comments Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2025
机构 * Santa Clara, CA, USA(美国圣克拉拉) ; Carnegie Mellon University(卡内基梅隆大学) ; Huazhong University of Science and Technology(华中科技大学) ; Nanjing University of Information Science and Technology(南京信息工程大学)
专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CL、eess.AS
Comments Accepted by Interspeech
Journal ref Proc. of Interspeech2025
机构 * Academia Sinica(台湾“中央研究院)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments Accepted to Interspeech 2025 Workshop
机构 * Tampere University, Tampere, Finland(塔尔库大学) ; University of Oxford, Oxford, UK(牛津大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Preprint version. The Version of Record is published in DAGM GCPR 2025 proceedings with Springer Lecture Notes in Computer Science (LNCS). Updated results and resources are available at the project page: https://saganet.notion.site
机构 * School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) ; Institute of Acoustics Chinese Academy of Science(中国科学院声学研究所)
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
Comments Submitted to ICASSP 2026
专题命中 音频语音多模态 :multimodal(abstract)
Comments 23 pages, 21 figures
机构 * Robotics Department, University of Michigan(密歇根大学机器人系)
专题命中 音频语音多模态 :multimodal(abstract)
Comments 8 pages, 7 figures
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; East China Normal University(华东师范大学) ; Soochow University(苏州大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI
机构 * Mohamed bin Zayed University of Artificial Intelligence(莫德赫·本·扎耶德人工智能大学) ; University of Central Florida(中央佛罗里达大学) ; Islamic University of Technology(伊斯兰技术大学) ; Air University(空军大学) ; ETH Zurich(苏黎世联邦理工学院) ; Technische Universität München(慕尼黑技术大学) ; National Institute of Informatics(国家信息研究所) ; Australian National University(澳大利亚国立大学) ; Linköping University(利尔贝里大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
机构 * UC Berkeley(伯克利大学) ; Stanford(斯坦福大学) ; UCL(伦敦大学学院) ; Virginia Tech(弗吉尼亚理工学院) ; Nvidia(英伟达公司)
专题命中 视频多模态 :multimodal(title,abstract)
机构 * Technical University of Munich(慕尼黑技术大学) ; ETH Zürich(苏黎世联邦理工学院) ; Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
专题命中 视频多模态 :cross-modal(title);分类 cs.CV
机构 * Peking University(北京大学) ; AI Geeks ; Australian Artificial Intelligence Institute(澳大利亚人工智能研究所)
专题命中 视频多模态 :MLLM(abstract);cross-modal(abstract);分类 cs.CV
机构 * Stanford University(斯坦福大学) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments ICCV Short Video Understanding Workshop Paper
机构 * NVIDIA ; MIT(麻省理工学院) ; HKU(香港大学) ; UC Berkeley(加州大学伯克利分校)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted by NeurIPS 2025. Code at https://github.com/NVlabs/Long-RL and model at https://huggingface.co/Efficient-Large-Model/LongVILA-R1-7B
机构 * IEEE CBMI 2025
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments IEEE CBMI 2025. This is the authors' accepted version. The final publication is available at https://ieeexplore.ieee.org/
机构 * Fantasyele
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments Under review
机构 * College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
Comments 18 pages,15 figures
Journal ref IEEE Transactions on Neural Networks and Learning Systems, pp. 1-15, 2025
机构 * Qualcomm AI Research(高通人工智能研究)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV
Comments Accepted to NeurIPS 2025
机构 * Peter Munk Cardiac Centre(彼得·默克心脏中心) ; Ted Rogers Centre for Heart Research(泰德·罗杰斯心脏病研究中心) ; University Health Network(大学健康网络) ; Joint Department of Medical Imaging(联合医学影像部门) ; University of Toronto(多伦多大学) ; Vector Institute(向量研究所)
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV
Comments ICCV 2025