Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments BMVC 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments BMVC 2025
机构 * Apple(苹果公司)
专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments TMLR August 2025
专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments 46 pages, 33 figures, Submitted to Advanced Engineering Informatics, under revision
机构 * Interactive Technologies Institute and NOVA LINCS Faculty of Exact Sciences and Engineering University of Madeira Portugal(互动技术研究所和NOVA LINCS精确科学与工程学院马德拉大学) ; Department of Informatics University of Bergen Norway(信息学院卑尔根大学挪威) ; Valencian Research Institute for Artificial Intelligence Universitat Politècnica de València Spain(瓦伦西亚人工智能研究机构瓦伦西亚理工大学西班牙) ; Leverhulme Centre for the Future of Intelligence and Valencian Research Institute for Artificial Intelligence Spain(未来智能中心和瓦伦西亚人工智能研究机构西班牙)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL
Comments 54 pages (42 pages of appendix). Accepted for publication at the ECAI 2025 conference
机构 * School of Artificial Intelligence and Software Engineering, Nanyang Normal University, Henan, China(人工智能与软件工程学院,南阳师范学院,河南) ; Institute for Artificial Intelligence, Peking University, Beijing, China(人工智能研究院,北京大学,北京) ; Collaborative Innovation Center of Intelligent Explosion-proof Equipment, Henan, China(智能防爆设备协同创新中心,河南)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted at the International Joint Conference on Artificial Intelligence (IJCAI 2025)
机构 * 1 Department of Artificial Intelligence, Sogang University, Seoul 04107, Republic of Korea 2 Department of Electronic Engineering, Sogang University, Seoul 04107, Republic of Korea 3 Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA 15213, USA 4 Mindslab Inc., Gyeonggi-do 13493, Republic of Korea 5 ICT Convergence Disaster/Safety Research Institute, Sogang University, Seoul 04107, Republic of Korea
专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to ICASSP 2024
机构 * Johns Hopkins University(约翰霍普金斯大学) ; Meta AI Research(Meta AI 研究)
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments EMNLP 2025 (Findings)
机构 * EPFL(苏黎世联邦理工学院) ; Idiap Research Institute(日内瓦研究所)
专题命中 音频语音多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM
Comments Accepted at ACM Multimedia 2025
专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)
机构 * Singapore Institute of Technology(新加坡理工学院) ; Duke Kunshan University(杜克-昆山大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments This paper has been accepted by APCIPA ASC 2025
机构 * Shanghai Jiao Tong University(上海交通大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM
机构 * Department of Computer Science and Engineering, Koç University(计算机科学与工程系,科克大学) ; Department of Psychology, Boğaziçi University(心理学系,博多伊大学) ; National Institute of Advanced Industrial Science and Technology (AIST), Intelligent Platforms Research Institute(国家先进工业科学与技术研究院(AIST),智能平台研究机构) ; Department of Psychology, Boğaziçi University University(心理学系,博多伊大学) ; Department of Computer Engineering, Hacettepe University(计算机工程系,哈切塞特佩大学) ; KUIS AI Center(KUIS人工智能中心)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV
Comments Accepted for publication in IEEE Transaction on Pattern Analysis and Machine Intelligence (IEEE TPAMI)
机构 * Singapore Institute of Technology(新加坡理工学院) ; Institute Of Acoustics, Chinese Academy Of Sciences(中国科学院声学研究所) ; Duke Kunshan University(杜克大学昆山分校)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
Comments The paper has been accepted by APCIPA ASC 2025
专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS
Comments 5 pages
机构 * Department of Electrical Engineering(电气工程系) ; Department of Semiconductor Engineering(半导体工程系) ; Center for Semiconductor Technology Convergence(半导体技术融合中心)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI
Journal ref Interspeech 2025
专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract)
Comments Accepted by KDD 2025
专题命中 视频多模态 :multimodal(title)
Comments 25 Pages and 5 Figures and a supplementary discussion as well
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; SpreeAI
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
Comments Project Page: https://immortalco.github.io/DressAndDance/
专题命中 视频多模态 :multimodal(abstract)
Comments Accepted to CoRL 2025; Github Page: https://long-vla.github.io
机构 * Northeastern University(东北大学) ; Memorial Sloan Kettering Cancer Center(纪念斯隆凯特琳癌症中心)
专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
机构 * Department of Computer Science, University of Virginia(大学计算机科学系) ; University of Virginia School of Medicine(弗吉尼亚大学医学院) ; Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) ; Biocomplexity Institute and Initiative, University of Virginia(大学生物复杂性研究所) ; Division of Infectious Diseases & International Health, University of Virginia School of Medicine(大学感染性疾病与国际卫生分会)
专题命中 跨模态检索 :multimodal(title,abstract)
机构 * Massachusetts Institute of Technology(麻省理工学院) ; Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所) ; Basis Research Institute(Basis研究机构) ; Eric and Wendy Schmidt Center(埃里克和温迪·施密特中心) ; BCAM (Basque Center for Applied Mathematics)(BCAM(巴斯克应用数学中心)) ; Ikerbasque (Basque Foundation for Science)(Ikerbasque(巴斯克科学基金会))
专题命中 跨模态检索 :multi-modal(abstract)
Comments 46 pages, 9 figures
机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队) ; Zhejiang University(浙江大学) ; Tsinghua University(清华大学)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Technical Report. Project Page: https://chenmingthu.github.io/milm/
机构 * MAUM AI Inc.(MAUM AI公司) ; Artificial Intelligence Graduate School UNIST(UNIST人工智能研究生学院)
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV
Comments Accepted to BMVC 2025
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV
Comments Code: github.com/showlab/Ego-PM
机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) ; Leibniz-Institut für Analytische Wissenschaften – ISAS – e.V.(莱比锡分析科学研究所(ISAS)) ; Department of Pathology, The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院病理科部) ; Institute of Pathology, University Hospital Essen(埃森大学医院病理科研究所) ; Academy for Multidisciplinary Studies, Capital Normal University(首都师范大学多学科研究学院)
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
机构 * Biometrics and Data Pattern Analytics Laboratory(生物信息与数据模式分析实验室) ; School of Engineering(工程学院)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments Published in Scientific Data (Nature). GitHub repository of the dataset at: https://github.com/BiDAlab/IMPROVE
Journal ref Scientific Data (2025) 12:1332
机构 * plaksha.edu.in(普拉克斯哈大学) ; research.iiit.ac.in(IIIT研究机构)
专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI
机构 * Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 多模态评测 :multimodal(abstract);分类 eess.AS