Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments BMVC 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments BMVC 2025
机构 * Apple(苹果公司)
专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments TMLR August 2025
专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments 46 pages, 33 figures, Submitted to Advanced Engineering Informatics, under revision
机构 * Interactive Technologies Institute and NOVA LINCS Faculty of Exact Sciences and Engineering University of Madeira Portugal(互动技术研究所和NOVA LINCS精确科学与工程学院马德拉大学) ; Department of Informatics University of Bergen Norway(信息学院卑尔根大学挪威) ; Valencian Research Institute for Artificial Intelligence Universitat Politècnica de València Spain(瓦伦西亚人工智能研究机构瓦伦西亚理工大学西班牙) ; Leverhulme Centre for the Future of Intelligence and Valencian Research Institute for Artificial Intelligence Spain(未来智能中心和瓦伦西亚人工智能研究机构西班牙)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL
Comments 54 pages (42 pages of appendix). Accepted for publication at the ECAI 2025 conference
机构 * School of Artificial Intelligence and Software Engineering, Nanyang Normal University, Henan, China(人工智能与软件工程学院,南阳师范学院,河南) ; Institute for Artificial Intelligence, Peking University, Beijing, China(人工智能研究院,北京大学,北京) ; Collaborative Innovation Center of Intelligent Explosion-proof Equipment, Henan, China(智能防爆设备协同创新中心,河南)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted at the International Joint Conference on Artificial Intelligence (IJCAI 2025)