arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-12 至 2025-08-12 共收录 100 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2508.07023 2025-08-12 cs.CV 85%

MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering

Jingwei Peng, Jiehao Chen, Mateo Alejandro Rojas, Meilin Zhang

机构 * Shaanxi University of Technology(陕西理工大学) Technological University of Peru(秘鲁技术大学)

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04377 2025-08-12 cs.CV cs.CL cs.MM 82%

Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions

Xiao Zhang, Johan Bos

机构 * University of Groningen(格罗宁根大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06623 2025-08-12 cs.CV 79%

ContextGuard-LVLM: Enhancing News Veracity through Fine-grained Cross-modal Contextual Consistency Verification

Sihan Ma, Qiming Wu, Ruotong Jiang, Frank Burns

机构 * Inner Mongolia University of Science & Technology(内蒙古科技大学) Federal University of Rio de Janeiro(里约热内卢联邦大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20756 2025-08-12 cs.CV cs.CL 73%

SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models

Zheng Liu, Hao Liang, Bozhou Li, Wentao Xiong, Chong Chen, Conghui He, Wentao Zhang, Bin Cui

机构 * Peking University Beijing China Huawei Technologies Ltd. Beijing China Shanghai AI Laboratory Shanghai China Peking University Center for Machine Learning Research Beijing China Peking University School of Computer Science \& Key Lab of High Confidence Software Technologies (MOE) Beijing China Peking University Huawei Technologies Ltd. Shanghai AI Laboratory

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07432 2025-08-12 cs.CV cs.AI 62%

Freeze and Reveal: Exposing Modality Bias in Vision-Language Models

Vivek Hruday Kavuri, Vysishtya Karanam, Venkata Jahnavi Venkamsetty, Kriti Madumadukala, Lakshmipathi Balaji Darur, Ponnurangam Kumaraguru

机构 * IIIT Hyderabad

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07318 2025-08-12 cs.CV 57%

RORPCap: Retrieval-based Objects and Relations Prompt for Image Captioning

Jinjing Gu, Tianbao Qin, Yuanyuan Pu, Zhengpeng Zhao

机构 * School of Information Science and Engineering(信息科学与工程学院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07260 2025-08-12 cs.CV 57%

Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM

Sihan Yang, Huitong Ji, Shaolin Lu, Jiayi Chen, Binxiao Xu, Ming Lu, Yuanxing Zhang, Wenhui Dong, Wentao Zhang

机构 * Peking University(北京大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06525 2025-08-12 cs.CV 57%

Large Language Models Facilitate Vision Reflection in Image Classification

Guoyuan An, JaeYoon Kim, SungEui Yoon

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2508.07608 2025-08-12 cs.MM cs.CV cs.SD eess.AS 85%

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition

Junxiao Xue, Xiaozhen Liu, Xuecheng Wu, Xinyi Yin, Danlei Huang, Fei Yu

机构 * Zhengzhou University(郑州大学) Zhejiang Lab(浙江实验室) Xi'an Jiaotong University(西安交通大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted by the ACM MM 2025 Workshop on SVC

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06902 2025-08-12 cs.CV 83%

eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos

Xuecheng Wu, Dingkang Yang, Danlei Huang, Xinyi Yin, Yifan Wang, Jia Zhang, Jiayu Nie, Liangyu Fu, Yang Liu, Junxiao Xue, Hadi Amirpour, Wei Zhou

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) College of Intelligent Robotics and Advanced Manufacturing, Fudan University & ByteDance(智能机器人与先进制造学院,复旦大学 & 字节跳动) School of Cyber Science and Engineering, Zhengzhou University(网络科学与工程学院,郑州大学) Institute of Advanced Technology, University of Science and Technology of China(先进技术研究院,中国科学技术大学) Inspur Electronic Information Industry Co., Ltd(Inspur电子信息产业有限公司) Department of Computer Science, The University of Toronto(计算机科学系,多伦多大学) Research Center for Space Computing System, Zhejiang Lab(空间计算系统研究中心,浙江实验室) Institute of Information Technology, University of Klagenfurt(信息技术学院,克雷格福特大学) School of Computer Science and Informatics, Cardiff University(计算机科学与信息学院,卡迪夫大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07337 2025-08-12 eess.AS cs.CV 81%

KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features

Ivan Kukanov, Jun Wah Ng

机构 * KLASS Engineering and Solutions Singapore(KLASS工程与解决方案新加坡)

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 cs.CV、eess.AS

Comments 7 pages, accepted to the 33rd ACM International Conference on Multimedia (MM'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08042 2025-08-12 cs.IR cs.AI 79%

Multi-modal Adaptive Mixture of Experts for Cold-start Recommendation

Van-Khang Nguyen, Duc-Hoang Pham, Huy-Son Nguyen, Cam-Van Thi Nguyen, Hoang-Quynh Le, Duc-Trong Le

机构 * VNU University of Engineering and Technology(越南工程大学) Delft University of Technology(代尔夫特理工大学)

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02516 2025-08-12 cs.CV 79%

Engagement Prediction of Short Videos with Large Multimodal Models

Wei Sun, Linhan Cao, Yuqin Cao, Weixia Zhang, Wen Wen, Kaiwei Zhang, Zijian Chen, Fangfang Lu, Xiongkuo Min, Guangtao Zhai

机构 * East China Normal University(华东师范大学) Shanghai Jiao Tong University(上海交通大学) City University of Hong Kong(香港城市大学) Shanghai University of Electric Power(上海电力大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments The proposed method achieves first place in the ICCV VQualA 2025 EVQA-SnapUGC Challenge on short-form video engagement prediction

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07625 2025-08-12 cs.CV 74%

A Trustworthy Method for Multimodal Emotion Recognition

Junxiao Xue, Xiaozhen Liu, Jie Wang, Xuecheng Wu, Bin Wu

机构 * Big Data Mining and Analytics, xxxxxxx 20xx, x(x): xxx-xxx(大数据挖掘与分析,xxxxx 20xx, x(x): xxx-xxx)

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV

Comments Accepted for publication in Big Data Mining and Analytics (BDMA), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07973 2025-08-12 cs.SD cs.CL eess.AS 62%

Joint Transcription of Acoustic Guitar Strumming Directions and Chords

Sebastian Murgul, Johannes Schimper, Michael Heizmann

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted to the 26th International Society for Music Information Retrieval Conference (ISMIR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 9 篇

2508.06939 2025-08-12 cs.AI cs.LG 79%

Intrinsic Explainability of Multimodal Learning for Crop Yield Prediction

Hiba Najjar, Deepak Pathak, Marlon Nuske, Andreas Dengel

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07401 2025-08-12 cs.CV 70%

LET-US: Long Event-Text Understanding of Scenes

Rui Chen, Xingyu Chen, Shaoan Wang, Shihan Kong, Junzhi Yu

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08035 2025-08-12 cs.CV cs.AI 62%

LVBench: An Extreme Long Video Understanding Benchmark

Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Xiaotao Gu, Shiyu Huang, Bin Xu, Yuxiao Dong, Ming Ding, Jie Tang

机构 * Zhipu AI(智谱AI) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07989 2025-08-12 cs.CV cs.HC 57%

The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility

Xiantao Zhang

机构 * Beihang University(北航大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 9 pages, 3 figures, 2 tables. Accepted at CV4A11y, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07626 2025-08-12 cs.CV cs.RO 57%

AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning

Dejie Yang, Zijing Zhao, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07312 2025-08-12 cs.CV 57%

MobileViCLIP: An Efficient Video-Text Model for Mobile Devices

Min Yang, Zihan Jia, Zhilin Dai, Sheng Guo, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) MyBank, Ant Group(蚂蚁集团MyBank) Shanghai AI Lab(上海AI实验室)

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07006 2025-08-12 eess.IV cs.CV 57%

Spatio-Temporal Conditional Diffusion Models for Forecasting Future Multiple Sclerosis Lesion Masks Conditioned on Treatments

Gian Mario Favero, Ge Ya Luo, Nima Fathi, Justin Szeto, Douglas L. Arnold, Brennan Nichyporuk, Chris Pal, Tal Arbel

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to MICCAI 2025 (LMID Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12524 2025-08-12 cs.CV cs.HC cs.LG eess.IV 57%

Inference-Time Gaze Refinement for Micro-Expression Recognition: Enhancing Event-Based Eye Tracking with Motion-Aware Post-Processing

Nuwan Bandara, Thivya Kandappu, Archan Misra

机构 * School of Computing(计算学院) Information Systems, Singapore Management University, Singapore(信息系统,新加坡管理大学,新加坡)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at 4DMR@IJCAI25: International IJCAI Workshop on 1st Challenge and Workshop for 4D Micro-Expression Recognition for Mind Reading, August 29, 2025, Guangzhou, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06892 2025-08-12 astro-ph.SR physics.space-ph 50%

Large Model Driven Solar Activity AI Forecaster: A Scalable Dual Data-Model Framework

Jingjing Wang, Pengyu Liang, Tingyu Wang, Ming Li, Yanmei Cui, Siwei Liu, Xin Huang, Xiang Li, Minghui Zhang, Yunshi Zeng, Zhu Cao, Jiekang Feng, Qinghua Hu, Bingxian Luo, Bing Cao

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2508.07666 2025-08-12 cs.MM 88%

Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts

Xianbing Zhao, Shengzun Yang, Buzhou Tang, Ronghuan Jiang

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.MM

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14904 2025-08-12 eess.IV cs.AI cs.CL cs.CV 85%

Large-vocabulary forensic pathological analyses via prototypical cross-modal contrastive learning

Chen Shen, Chunfeng Lian, Wanqing Zhang, Fan Wang, Jianhua Zhang, Shuanliang Fan, Xin Wei, Gongji Wang, Kehan Li, Hongshu Mu, Hao Wu, Xinggong Liang, Jianhua Ma, Zhenyuan Wang

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 28 pages, 6 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07671 2025-08-12 cs.AI cs.CY cs.HC cs.MA stat.AP 57%

EMPATHIA: Multi-Faceted Human-AI Collaboration for Refugee Integration

Mohamed Rayan Barhdadi, Mehmet Tuncel, Erchin Serpedin, Hasan Kurban

机构 * Texas A&M University(德克萨斯A&M大学) Istanbul Technical University(伊斯坦布尔技术大学) Hamad Bin Khalifa University(哈利法大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 19 pages, 3 figures (plus 6 figures in supplementary), 2 tables, 1 algorithm. Submitted to NeurIPS 2025 Creative AI Track: Humanity

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16919 2025-08-12 cs.CV 57%

TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

Xuying Zhang, Yutong Liu, Yangguang Li, Renrui Zhang, Yufei Liu, Kai Wang, Wanli Ouyang, Zhiwei Xiong, Peng Gao, Qibin Hou, Ming-Ming Cheng

机构 * VCIP, CS, Nankai University(南开大学计算机科学与技术学院) NKIARI, Shenzhen Futian(深圳未来科技研究院) USTC(University of Science and Technology of China) CUHK MMLab(香港中文大学MMLab) VAST(中国科学院自动化研究所) Shanghai AI Lab(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Project page: https://github.com/HVision-NKU/TAR3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06691 2025-08-12 cond-mat.mtrl-sci cs.LG 50%

Role of Large Language Models and Retrieval-Augmented Generation for Accelerating Crystalline Material Discovery: A Systematic Review

Agada Joseph Oche, Arpan Biswas

机构 * Bredesen Center for Interdisciplinary Research, University of Tennessee, Knoxville, USA(跨学科研究中心,田纳西大学,诺克斯维尔,美国) University of Tennessee-Oak Ridge Innovation Institute, University of Tennessee, Knoxville, USA(田纳西大学橡树岭创新研究所,田纳西大学,诺克斯维尔,美国)

专题命中 跨模态检索 :multi-modal(abstract)

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 16 篇

2503.06134 2025-08-12 cs.CV 83%

X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation

Jian Ma, Qirong Peng, Xu Guo, Chen Chen, Haonan Lu, Zhenyu Yang

机构 * OPPO AI Center(OPPO人工智能中心) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏