arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-18 至 2025-09-18 共收录 44 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 2 篇

2408.08872 2025-09-18 cs.CV cs.AI cs.CL 85%

xGen-MM (BLIP-3): A Family of Open Large Multimodal Models

Le Xue, Manli Shu, Anas Awadalla, Jun Wang, An Yan, Senthil Purushwalkam, Honglu Zhou, Viraj Prabhu, Yutong Dai, Michael S Ryoo, Shrikant Kendre, Jieyu Zhang, Shaoyen Tseng, Gustavo A Lujan-Moreno, Matthew L Olson, Musashi Hinck, David Cobbley, Vasudev Lal, Can Qin, Shu Zhang, Chia-Chih Chen, Ning Yu, Juntao Tan, Tulika Manoj Awalgaonkar, Shelby Heinecke, Huan Wang, Yejin Choi, Ludwig Schmidt, Zeyuan Chen, Silvio Savarese, Juan Carlos Niebles, Caiming Xiong, Ran Xu

机构 * Salesforce AI Research(Salesforce AI研究院) Intel Labs(英特尔实验室) University of Washington(华盛顿大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23759 2025-09-18 cs.CL cs.AI cs.CV cs.LG 67%

Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint

Heekyung Lee, Jiaxin Ge, Tsung-Han Wu, Minwoo Kang, Trevor Darrell, David M. Chan

机构 * POSTECH University of California, Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2509.14097 2025-09-18 cs.CV cs.MM 88%

Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing

Yaru Chen, Ruohao Guo, Liting Gao, Yang Xiang, Qingyu Luo, Zhenbo Li, Wenwu Wang

专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09595 2025-09-18 cs.CV 85%

Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis

Yikang Ding, Jiwen Liu, Wenyuan Zhang, Zekun Wang, Wentao Hu, Liyuan Cui, Mingming Lao, Yingchao Shao, Hui Liu, Xiaohan Li, Ming Chen, Xiaoqiang Liu, Yu-Shen Liu, Pengfei Wan

机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);audio-visual(abstract);分类 cs.CV

Comments Technical Report. Project Page: https://klingavatar.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13395 2025-09-18 eess.AS cs.AI cs.CL cs.LG cs.MM 83%

TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models

Haolong Zheng, Yekaterina Yegorova, Mark Hasegawa-Johnson

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15871 2025-09-18 cs.CY cs.AI cs.CL 62%

A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcare

Manar Aljohani, Jun Hou, Sindhura Kommu, Xuan Wang

机构 * Virginia Tech(维吉尼亚理工大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14023 2025-09-18 cs.CL cs.HC 57%

Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality

Sami Ul Haq, Sheila Castilho, Yvette Graham

机构 * ADAPT Centre(ADAPT中心) Dublin City University (DCU)(都柏林城市大学) Trinity College Dublin (TCD)(三一学院都柏林)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted at WMT2025 (ENNLP) for oral presented

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10432 2025-09-18 q-bio.OT cs.AI 57%

Standards in the Preparation of Biomedical Research Metadata: A Bridge2AI Perspective

Harry Caufield, Satrajit Ghosh, Sek Wong Kong, Jillian Parker, Nathan Sheffield, Bhavesh Patel, Andrew Williams, Timothy Clark, Monica C. Munoz-Torres

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03813 2025-09-18 cs.HC 50%

Talk to the Wall: The Role of Speech Interaction in Collaborative Visual Analytics

Gabriela Molina León, Anastasia Bezerianos, Olivier Gladin, Petra Isenberg

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages, 6 figures, to appear in IEEE TVCG (VIS 2024); correct figure

Journal ref IEEE Transactions on Visualization and Computer Graphics, 31(1), 2025, 941-951

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2509.13515 2025-09-18 cs.CV 79%

Multimodal Hate Detection Using Dual-Stream Graph Neural Networks

Jiangbei Yue, Shuonan Yang, Tailin Chen, Jianbo Jiao, Zeyu Fu

机构 * Multimodal Intelligence Lab, Department of Computer Science University of Exeter Exeter, UK(埃克塞特大学计算机科学系多模态智能实验室) School of Computer Science University of Birmingham Birmingham, UK(伯明翰大学计算机科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15864 2025-09-18 cs.RO 78%

FlowAct: A Proactive Multimodal Human-robot Interaction System with Continuous Flow of Perception and Modular Action Sub-systems

Timothée Dhaussy, Bassam Jabaian, Fabrice Lefèvre

机构 * Laboratoire Informatique d'Avignon, Avignon University, France(阿维尼翁信息实验室,阿维尼翁大学,法国)

专题命中 视频多模态 :multimodal(title,abstract)

Comments Paper accepted at ICPRAM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11915 2025-09-18 cs.SD cs.CV cs.LG cs.MM eess.AS 67%

Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound

Junwon Lee, Jaekwon Im, Dabin Kim, Juhan Nam

机构 * Graduate School of AI, KAIST(韩国国立庆熙大学人工智能研究生院) Graduate School of CT, KAIST(韩国国立庆熙大学CT研究生院)

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13722 2025-09-18 cs.CV cs.AI 62%

Mitigating Query Selection Bias in Referring Video Object Segmentation

Dingwei Zhang, Dong Zhang, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) The Hong Kong University of Science and Technology(香港科学大学) Nanjing Forestry University(南京林业大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 7 篇

2509.02962 2025-09-18 cs.CV 83%

Resilient Multimodal Industrial Surface Defect Detection with Uncertain Sensors Availability

Shuai Jiang, Yunfeng Ma, Jingyu Zhou, Yuan Bian, Yaonan Wang, Min Liu

机构 * School of Artificial Intelligence and Robotics(人工智能与机器人学院) National Engineering Research Center for Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心) Hunan University(湖南大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE/ASME Transactions on Mechatronics

Journal ref IEEE/ASME Transactions on Mechatronics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14142 2025-09-18 cs.CV 80%

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

Peng Xu, Shengwu Xiong, Jiajun Zhang, Yaxiong Chen, Bowen Zhou, Chen Change Loy, David A. Clifton, Kyoung Mu Lee, Luc Van Gool, Ruiming He, Ruilin Yao, Xinwei Long, Jirui Huang, Kai Tian, Sa Yang, Yihua Shao, Jin Feng, Yue Zhong, Jiakai Zhou, Cheng Tang, Tianyu Zou, Yifang Zhang, Junming Liang, Guoyou Li, Zhaoxiang Wang, Qiang Zhou, Yichen Zhao, Shili Xiong, Hyeongjin Nam, Jaerin Lee, Jaeyoung Chung, JoonKyu Park, Junghun Oh, Kanggeon Lee, Wooseok Lee, Juneyoung Ro, Turghun Osman, Can Hu, Chaoyang Liao, Cheng Chen, Chengcheng Han, Chenhao Qiu, Chong Peng, Cong Xu, Dailin Li, Feiyu Wang, Feng Gao, Guibo Zhu, Guopeng Tang, Haibo Lu, Han Fang, Han Qi, Hanxiao Wu, Haobo Cheng, Hongbo Sun, Hongyao Chen, Huayong Hu, Hui Li, Jiaheng Ma, Jiang Yu, Jianing Wang, Jie Yang, Jing He, Jinglin Zhou, Jingxuan Li, Josef Kittler, Lihao Zheng, Linnan Zhao, Mengxi Jia, Muyang Yan, Nguyen Thanh Thien, Pu Luo, Qi Li, Shien Song, Shijie Dong, Shuai Shao, Shutao Li, Taofeng Xue, Tianyang Xu, Tianyi Gao, Tingting Li, Wei Zhang, Weiyang Su, Xiaodong Dong, Xiao-Jun Wu, Xiaopeng Zhou, Xin Chen, Xin Wei, Xinyi You, Xudong Kang, Xujie Zhou, Xusheng Liu, Yanan Wang, Yanbin Huang, Yang Liu, Yang Yang, Yanglin Deng, Yashu Kang, Ye Yuan, Yi Wen, Yicen Tian, Yilin Tao, Yin Tang, Yipeng Lin, Yiqing Wang, Yiting Xi, Yongkang Yu, Yumei Li, Yuxin Qin, Yuying Chen, Yuzhe Cen, Zhaofan Zou, Zhaohong Liu, Zhehao Shen, Zhenglin Du, Zhengyang Li, Zhenni Huang, Zhenwei Shao, Zhilong Song, Zhiyong Feng, Zhiyu Wang, Zhou Yu, Ziang Li, Zihan Zhai, Zijian Zhang, Ziyang Peng, Ziyun Xiao, Zongshu Li

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025 MARS2 Workshop and Challenge "Multimodal Reasoning and Slow Thinking in the Large Model Era: Towards System 2 and Beyond''

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13474 2025-09-18 cs.CV 79%

Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization

Yujia Lin, Nicholas Evans

机构 * Dali University(大理大学) Bandırma Onyedi Eylül University(巴尔迪马十一点大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13888 2025-09-18 cs.CL cs.AI cs.IR 76%

Combating Biomedical Misinformation through Multi-modal Claim Detection and Evidence-based Verification

Mariano Barone, Antonio Romano, Giuseppe Riccio, Marco Postiglione, Vincenzo Moscato

机构 * University of Naples Federico II(那不勒斯费迪里奇二世大学) Northwestern University(西北大学)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL、cs.AI

Journal ref SIGIR '25: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05427 2025-09-18 cs.MM 57%

Reply with Sticker: New Dataset and Model for Sticker Retrieval

Bin Liang, Bingbing Wang, Zhixin Bai, Qiwei Lang, Mingwei Sun, Kaiheng Hou, Lanjun Zhou, Ruifeng Xu, Kam-Fai Wong

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.MM

Journal ref Liang B, Wang B, Bai Z, et al. Reply with Sticker: New Dataset and Model for Sticker Retrieval[J]. IEEE Transactions on Audio, Speech and Language Processing, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13326 2025-09-18 cs.HC cs.LG 50%

LLM Chatbot-Creation Approaches

Hemil Mehta, Tanvi Raut, Kohav Yadav, Edward F. Gehringer

专题命中 跨模态检索 :multimodal(abstract)

Comments Forthcoming in Frontiers in Education (FIE 2025), Nashville, Tennessee, USA, Nov 2-5, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12824 2025-09-18 cs.IR 50%

DiffHash: Text-Guided Targeted Attack via Diffusion Models against Deep Hashing Image Retrieval

Zechao Liu, Zheng Zhou, Xiangkun Chen, Tao Liang, Dapeng Lang

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 6 篇

2509.13642 2025-09-18 cs.LG cs.CV 83%

LLM-I: LLMs are Naturally Interleaved Multimodal Creators

Zirun Guo, Feng Zhang, Kai Jia, Tao Jin

机构 * Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title);MLLM(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01086 2025-09-18 cs.CV cs.AI 81%

DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing

Xiaolong Wang, Zhi-Qi Cheng, Jue Wang, Xiaojiang Peng

机构 * Shenzhen Technology University(深圳科技大学) University of Washington(华盛顿大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 13 pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13227 2025-09-18 math.OC cs.AI cs.SY eess.SY 74%

Rich Vehicle Routing Problem in Disaster Management enabling Temporally-causal Transhipments across Multi-Modal Transportation Network

Santanu Banerjee, Goutam Sen, Siddhartha Mukhopadhyay

机构 * Department of Industrial and Systems Engineering (ISE), Indian Institute of Technology (IIT) Kharagpur(工业与系统工程系,印度理工学院Kharagpur分校)

专题命中 多模态生成 :multi-modal(title);分类 cs.AI

Comments Major changes in version II: 1) Supplementary is now a separate document, 2) Algorithm steps have been updated with pseudocode in the Heuristic, 3) Explanation of the MILP formulation construction is further detailed in a supplementary section

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14837 2025-09-18 cs.RO cs.LG 71%

Learning Multimodal Attention for Manipulating Deformable Objects with Changing States

Namiko Saito, Mayu Tatsumi, Ayuna Kubo, Kanata Suzuki, Hiroshi Ito, Shigeki Sugano, Tetsuya Ogata

机构 * Future Robotics Organization, Waseda University(早稻田大学未来机器人组织) Microsoft Research Asia(微软亚洲研究院) Department of Modern Mechanical Engineering, Waseda University(早稻田大学现代机械工程系) Artificial Intelligence Laboratories, Fujitsu Limited(Fujitsu 人工智能实验室) Center for Technology Innovation - Controls and Robotics, Research & Development Group, Hitachi, Ltd.(富士通技术研发集团技术创新中心 - 控制与机器人) Faculty of Science and Engineering, Waseda University(早稻田大学工学部) National Institute of Advanced Science and Technology(国家先进科学研究院)

专题命中 多模态生成 :multimodal(title)

Comments Humanoids2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13760 2025-09-18 cs.CV 57%

Iterative Prompt Refinement for Safer Text-to-Image Generation

Jinwoo Jeon, JunHyeok Oh, Hayeong Lee, Byung-Jun Lee

机构 * Korea University(韩国大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13696 2025-09-18 cs.CL 57%

Integrating Text and Time-Series into (Large) Language Models to Predict Medical Outcomes

Iyadh Ben Cheikh Larbi, Ajay Madhavan Ravichandran, Aljoscha Burchardt, Roland Roller

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Technical University Berlin(柏林技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments Presented and published at BioCreative IX

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 4 篇

2509.13773 2025-09-18 cs.AI cs.IR 83%

MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation

Zhipeng Bian, Jieming Zhu, Xuyang Xie, Quanyu Dai, Zhou Zhao, Zhenhua Dong

机构 * Shenzhen University(深圳大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Zhejiang University(浙江大学)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.AI

Comments Published in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), ACL 2025. Official version: https://doi.org/10.18653/v1/2025.acl-industry.103

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) ACL 2025 1457-1465

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16405 2025-09-18 cs.MM 83%

EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment

Lancheng Gao, Ziheng Jia, Yunhao Zeng, Wei Sun, Yiming Zhang, Wei Zhou, Guangtao Zhai, Xiongkuo Min

专题命中 多模态评测 :multi-modal(title,abstract);MLLM(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12248 2025-09-18 cs.CV cs.AI cs.CL 82%

Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics

Yuriel Ryan, Rui Yang Tan, Kenny Tsu Wei Choo, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 27 pages, 8 figures, EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13692 2025-09-18 cs.RO 78%

HGACNet: Hierarchical Graph Attention Network for Cross-Modal Point Cloud Completion

Yadan Zeng, Jiadong Zhou, Xiaohan Li, I-Ming Chen

机构 * Robotics Research Centre of the School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore(南洋理工大学机械与航空航天工程学院机器人研究中心) College of Information and Control Engineering, Xi’an University of Architecture and Technology, Xi’an, China(西安建筑科技大学信息与控制工程学院)

专题命中 多模态评测 :cross-modal(title,abstract)

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏