arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-09 至 2025-10-09 共收录 44 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2505.02152 2025-10-09 cs.RO 82%

Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions

Cunxin Fan, Xiaosong Jia, Yihang Sun, Yixiao Wang, Jianglan Wei, Ziyang Gong, Xiangyu Zhao, Masayoshi Tomizuka, Xue Yang, Junchi Yan, Mingyu Ding

机构 * Shanghai Jiao Tong University(上海交通大学) UC Berkeley(伯克利大学) UNC, Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03363 2025-10-09 cs.CV cs.AI eess.IV 62%

Unified Unsupervised Anomaly Detection via Matching Cost Filtering

Zhe Zhang, Mingxiu Cai, Gaochang Wu, Jing Zhang, Lingqiao Liu, Dacheng Tao, Tianyou Chai, Xiatian Zhu

机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang, China(合成过程工业综合自动化国家重点实验室,东北大学,沈阳,中国) University of Surrey(Surrey大学) School of Computer Science, Wuhan University(武汉大学计算机学院) School of Computer Science, The University of Adelaide(阿德莱德大学计算机学院) College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) Surrey Institute for People-Centred Artificial Intelligence, and Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey人本人工智能研究所,以及视觉、语音和信号处理中心,Surrey大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 63 pages (main paper and supplementary material), 39 figures, 58 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18269 2025-10-09 cs.CV cs.AI 62%

MAMS: Model-Agnostic Module Selection Framework for Video Captioning

Sangho Lee, Il Yong Chun, Hogun Park

机构 * Sangho Lee 1,2(Sangho Lee 教授) Il Yong Chun 1,3(Il Yong Chun 教授) Hogun Park 1(Hogun Park 教授)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to the AAAI 2025 Main Technical Track. This is an extended version of the original submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07098 2025-10-09 cs.CL 57%

TALENT: Table VQA via Augmented Language-Enhanced Natural-text Transcription

Guo Yutong, Wanying Wang, Yue Wu, Zichen Miao, Haoyu Wang

机构 * Johns Hopkins University(约翰霍普金斯大学) Purdue University(普渡大学) University at Albany(阿尔巴尼大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2403.16276 2025-10-09 cs.CV cs.AI 84%

Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding

Yolo Yunlong Tang, Daiki Shimada, Jing Bi, Mingqian Feng, Hang Hua, Chenliang Xu

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07839 2025-10-09 cs.RO 78%

Touch Speaks, Sound Feels: A Multimodal Approach to Affective and Social Touch from Robots to Humans

Qiaoqiao Ren, Tony Belpaeme

机构 * Faculty of Engineering and Architecture, IDLab-AIRO, Ghent University – imec, Technologiepark 126, 9052 Gent, Belgium(工程与建筑学院,IDLab-AIRO,根特大学–imec,Technologiepark 126,9052 Gent,比利时)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06872 2025-10-09 cs.HC 78%

Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz

Frederic Gmeiner, Kenneth Holstein, Nikolas Martelaro

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 18 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07293 2025-10-09 cs.SD cs.AI cs.CL eess.AS 67%

AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs

Peize He, Zichen Wen, Yubo Wang, Yuxuan Wang, Xiaoqian Liu, Jiajie Huang, Zehui Lei, Zhuangcheng Gu, Xiangqi Jin, Jiabing Yang, Kai Li, Zhifei Liu, Weijia Li, Cunxiang Wang, Conghui He, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Northeastern University(东北大学) Carnegie Mellon University(卡内基梅隆大学) University of Chinese Academy of Sciences(中国科学院大学) Tsinghua University(清华大学) Sun Yat-sen University(中山大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments 26 pages, 23 figures, the code is available at \url{https://github.com/DabDans/AudioMarathon}

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16765 2025-10-09 cs.CL cs.AI cs.SD eess.AS 67%

The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology

Fagun Patel, Duc Q. Nguyen, Sang T. Truong, Jody Vaynshtok, Sanmi Koyejo, Nick Haber

机构 * Stanford University(斯坦福大学) National University of Singapore(新加坡国立大学) Sound Speech and Hearing Clinic(语音与听力诊所)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments EMNLP 2025 Oral Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07299 2025-10-09 eess.AS cs.SD 57%

Comparison of Speech Tasks in Human Expert and Machine Detection of Parkinson's Disease

Peter Plantinga, Roozbeh Sattari, Karine Marcotte, Carla Di Gironimo, Madeleine Sharp, Liziane Bouvier, Maiya Geddes, Ingrid Verduyckt, Étienne de Villers-Sidani, Mirco Ravanelli, Denise Klein

机构 * McGill University(麦吉尔大学) CRBLM Mila Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) Nouvelle Voix(新声音) Montreal Neurological Institute(蒙特利尔神经科学研究所) Concordia University(Concordia大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted to SMASH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 6 篇

2404.12353 2025-10-09 cs.CV cs.AI 84%

V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning

Hang Hua, Yolo Yunlong Tang, Chenliang Xu, Jiebo Luo

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12164 2025-10-09 cs.CV cs.AI cs.MM 82%

Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward

Yolo Yunlong Tang, Siting Xu, Teng Wang, Qin Lin, Qinglin Lu, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) Tencent Inc.(腾讯公司)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by ACCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12009 2025-10-09 cs.CV 70%

CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion

Yolo Yunlong Tang, Gen Zhan, Li Yang, Yiting Liao, Chenliang Xu

机构 * ByteDance(字节跳动)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06235 2025-10-09 eess.IV cs.AI cs.CV q-bio.NC 62%

Stacked Regression using Off-the-shelf, Stimulus-tuned and Fine-tuned Neural Networks for Predicting fMRI Brain Responses to Movies (Algonauts 2025 Report)

Robert Scholz, Kunal Bagga, Christine Ahrends, Carlo Alberto Barbano

机构 * Université Paris Cité(巴黎Cité大学) University of Oxford(牛津大学) University of Turin(都灵大学) Universität Leipzig(莱比锡大学) Max Planck School of Cognition(马克斯·普朗克认知科学学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06973 2025-10-09 cs.CV 57%

Addressing the ID-Matching Challenge in Long Video Captioning

Zhantao Yang, Huangji Wang, Ruili Feng, Han Zhang, Yuting Hu, Shangwen Zhu, Junyan Li, Yu Liu, Fan Cheng

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba group(阿里巴巴集团)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06657 2025-10-09 cs.IR 50%

LLM-Powered Nuanced Video Attribute Annotation for Enhanced Recommendations

Boyuan Long, Yueqi Wang, Hiloni Mehta, Mick Zomnir, Omkar Pathak, Changping Meng, Ruolin Jia, Yajun Peng, Dapeng Hong, Xia Wu, Mingyan Gao, Onkar Dalal, Ningren Han

专题命中 视频多模态 :multimodal(abstract)

Comments RecSys 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2510.06888 2025-10-09 cs.IR cs.AI 83%

M3Retrieve: Benchmarking Multimodal Retrieval for Medicine

Arkadeep Acharya, Akash Ghosh, Pradeepika Verma, Kitsuchart Pasupa, Sriparna Saha, Priti Singh

机构 * Indian Institute of Technology Patna(印度理工学院帕纳瓦分校) King Mongkut’s Institute of Technology Ladkrabang(拉差丹awan技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments EMNLP Mains 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03268 2025-10-09 cs.LG cs.AI 79%

Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment

Lingjie Yi, Raphael Douady, Chao Chen

机构 * Stony Brook University(史坦尼·布鲁克大学) University Paris 1 Pantheon-Sorbonne(巴黎第一大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11557 2025-10-09 cs.CR cs.AI 74%

AC-LoRA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs

Lara Magdalena Lazier, Aritra Dhar, Vasilije Stambolic, Lukas Cavigelli

专题命中 跨模态检索 :multi-modal(title);分类 cs.AI

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20444 2025-10-09 cs.LG cs.CV 57%

HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models

Haoran Li, Yingjie Qin, Baoyuan Ou, Lai Xu, Ruiwen Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Xiaohongshu Inc.(小红书公司)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 3 篇

2510.06679 2025-10-09 cs.CV 79%

DreamOmni2: Multimodal Instruction-based Editing and Generation

Bin Xia, Bohao Peng, Yuechen Zhang, Junjia Huang, Jiyang Liu, Jingyao Li, Haoru Tan, Sitong Wu, Chengyao Wang, Yitong Wang, Xinglong Wu, Bei Yu, Jiaya Jia

机构 * CUHK(香港中文大学) HKUST(香港科技大学) HKU(香港大学) ByteDance Inc(字节跳动公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06308 2025-10-09 cs.CV 79%

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

Yi Xin, Qi Qin, Siqi Luo, Kaiwen Zhu, Juncheng Yan, Yan Tai, Jiayi Lei, Yuewen Cao, Keqi Wang, Yibin Wang, Jinbin Bai, Qian Yu, Dengyang Jiang, Yuandong Pu, Haoxing Chen, Le Zhuo, Junjun He, Gen Luo, Tianbin Li, Ming Hu, Jin Ye, Shenglong Ye, Bo Zhang, Chang Xu, Wenhai Wang, Hongsheng Li, Guangtao Zhai, Tianfan Xue, Bin Fu, Xiaohong Liu, Yu Qiao, Yihao Liu

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Nanjing University(南京大学) The University of Sydney(悉尼大学) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 33 pages, 13 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07627 2025-10-09 cs.CE q-bio.BM 50%

LSMTCR: A Scalable Multi-Architecture Model for Epitope-Specific T Cell Receptor de novo Design

Ruihao Zhang, Xiao Liu

专题命中 多模态生成 :cross-modal(abstract)

Comments 13 main pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 14 篇

2412.09278 2025-10-09 cs.CV cs.AI 84%

Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine

Xiaoshuang Huang, Lingdong Shen, Jia Liu, Fangxin Shang, Hongxiang Li, Haifeng Huang, Yehui Yang

机构 * Baidu Inc.(百度公司)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18562 2025-10-09 cs.CL cs.AI 84%

GIIFT: Graph-guided Inductive Image-free Multimodal Machine Translation

Jiafeng Xiong, Yuting Zhao

机构 * Department of Computer Science University of Manchester(曼彻斯特大学计算机科学系) Department of Advanced Information Technology Kyushu University(九州大学先进信息技术系)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted as an oral presentation at the EMNLP 2025 Workshop on Machine Translation (WMT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14146 2025-10-09 cs.CL 79%

MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation

Xian Gao, Jiacheng Ruan, Zongyun Zhang, Jingsheng Gao, Ting Liu, Yuzhuo Fu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06252 2025-10-09 q-bio.NC cs.AI cs.LG 79%

Dream2Image : An Open Multimodal EEG Dataset for Decoding and Visualizing Dreams with Artificial Intelligence

Yann Bellec

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 7 Pages, 3 Figures, The Dream2Image dataset is openly available on Hugging Face at: https://huggingface.co/datasets/opsecsystems/Dream2Image

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19017 2025-10-09 cs.CL 79%

Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models

Bin Zhu, Yinxuan Gui, Huiyan Qi, Jingjing Chen, Chong-Wah Ngo, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学) Fudan University(复旦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Project website: https://yxg1005.github.io/GaslightingNegationAttacks/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06633 2025-10-09 cs.RO 78%

Assist-As-Needed: Adaptive Multimodal Robotic Assistance for Medication Management in Dementia Care

Kruthika Gangaraju, Tanmayi Inaparthy, Jiaqi Yang, Yihao Zheng, Fengpei Yuan

机构 * Worcester Polytechnic Institute(沃斯特理工学院)

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06743 2025-10-09 cs.CV cs.AI cs.CL 67%

Evaluating LLMs for Historical Document OCR: A Methodological Framework for Digital Humanities

Maria Levchenko

机构 * Italian Institute of Germanic Studies (IISG)(意大利德语研究学院) University of Bologna(博洛尼亚大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The First Workshop on Natural Language Processing and Language Models for Digital Humanities (LM4DH 2025). RANLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏