arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-07 至 2025-10-07 共收录 98 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 14 篇

2405.14715 2025-10-07 cs.CV cs.AI 84%

Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models

Young Kyun Jang, Ser-nam Lim

机构 * Google DeepMind(谷歌DeepMind) University of Central Florida(中央佛罗里达大学)

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02608 2025-10-07 cs.AI 83%

Mitigating Modal Imbalance in Multimodal Reasoning

Chen Henry Wu, Neil Kale, Aditi Raghunathan

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 10 pages, 10 figures, CoLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04145 2025-10-07 cs.CV cs.CL cs.IR 81%

Automating construction safety inspections using a multi-modal vision-language RAG framework

Chenxin Wang, Elyas Asadi Shamsabadi, Zhaohui Chen, Luming Shen, Alireza Ahmadian Fard Fini, Daniel Dias-da-Costa

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments 33 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03295 2025-10-07 cs.CV cs.CL cs.LG 81%

Multimodal Arabic Captioning with Interpretable Visual Concept Integration

Passant Elchafei, Amany Fashwan

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18743 2025-10-07 cs.CV 79%

SAR-TEXT: A Large-Scale SAR Image-Text Dataset Built with SAR-Narrator and A Progressive Learning Strategy for Downstream Tasks

Yiguo He, Xinjun Cheng, Junjie Zhu, Chunping Qiu, Jun Wang, Xichuan Zhang, Qiangjuan Huang, Ke Yang

机构 * Intelligent Game and Decision Lab(智能游戏与决策实验室)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments IEEE Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16856 2025-10-07 cs.CV cs.AI 74%

SIA: Enhancing Safety via Intent Awareness for Vision-Language Models

Youngjin Na, Sangheon Jeong, Youngwan Lee, Jian Lee, Dawoon Jeong, Youngman Kim

机构 * VLM Safety LAB, MODULABS(视觉语言模型安全实验室,MODULABS) ETRI(电子技术研究院) KAIST(韩国科学技术院)

专题命中 图文多模态 :multimodal(abstract,comments);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted to Safe and Trustworthy Multimodal AI Systems(SafeMM-AI) Workshop at ICCV2025, Non-archival track

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16965 2025-10-07 cs.CL cs.CV 73%

Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning

Zhe Hu, Jing Li, Zhongzhu Pu, Hou Pong Chan, Yu Yin

机构 * The Hong Kong Polytechnic University(香港理工大学) Tsinghua University(清华大学) InspireOmni AI Alibaba Group(阿里巴巴集团) Case Western Reserve University(凯斯西储大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10610 2025-10-07 cs.CV cs.CL 62%

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Zhaowei Wang, Wenhao Yu, Xiyu Ren, Jipeng Zhang, Yu Zhao, Rohit Saxena, Liang Cheng, Ginny Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman

机构 * CSE Department, HKUST(香港科技大学计算机科学与工程系) Tencent AI Seattle Lab(腾讯AI西雅图实验室) University of Edinburgh(爱丁堡大学) NVIDIA AI Technology Center (NVAITC), NVIDIA, Santa Clara, USA(英伟达圣克拉拉人工智能技术中心)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted as a spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03483 2025-10-07 cs.CV cs.AI 62%

DuPLUS: Dual-Prompt Vision-Language Framework for Universal Medical Image Segmentation and Prognosis

Numan Saeed, Tausifa Jan Saleem, Fadillah Maani, Muhammad Ridzuan, Hu Wang, Mohammad Yaqub

机构 * Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(计算机视觉系,Mohamed bin Zayed人工智能大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03441 2025-10-07 cs.CV cs.AI cs.LG 62%

Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning

Chashi Mahiul Islam, Oteo Mamo, Samuel Jacob Chacko, Xiuwen Liu, Weikuan Yu

机构 * Florida State University(佛罗里达州立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22229 2025-10-07 cs.CV 57%

A Tale of Two Experts: Cooperative Learning for Source-Free Unsupervised Domain Adaptation

Jiaping Yu, Muli Yang, Jiapeng Ji, Jiexi Yan, Cheng Deng

机构 * School of Electronic Engineering(电子工程学院) Xidian University(西安电子科技大学) School of Computer Science(计算机科学学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03858 2025-10-07 cs.CV 57%

Cross-View Open-Vocabulary Object Detection in Aerial Imagery

Jyoti Kini, Rohit Gupta, Mubarak Shah

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18857 2025-10-07 cs.CV cs.LG 57%

Probabilistic Language-Image Pre-Training

Sanghyuk Chun, Wonjae Kim, Song Park, Sangdoo Yun

机构 * NAVER AI Lab(NAVER AI实验室)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Code: https://github.com/naver-ai/prolip HuggingFace Hub: https://huggingface.co/collections/SanghyukChun/prolip-6712595dfc87fd8597350291 33 pages, 4.5 MB; LongProLIP paper: arXiv:2503.08048; Multiplicity paper for more background: arxiv.org:2505.19614; v4: fix typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04710 2025-10-07 cs.LG 50%

ViTs: Teaching Machines to See Time Series Anomalies Like Human Experts

Zexin Wang, Changhua Pei, Yang Liu, Hengyue Jiang, Quan Zhou, Haotian Si, Hang Cui, Jianhui Li, Gaogang Xie, Jingjing Li, Dan Pei

机构 * Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州先进研究所) Tsinghua University(清华大学)

专题命中 图文多模态 :image-text(abstract)

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 8 篇

2501.04686 2025-10-07 cs.CL cs.AI cs.LG 84%

Unlocking Multimodal Mathematical Reasoning via Process Reward Model

Ruilin Luo, Zhuofan Zheng, Yifan Wang, Xinzhe Ni, Zicheng Lin, Songtao Jiang, Yiyao Yu, Chufan Shi, Lei Wang, Ruihang Chu, Jin Zeng, Yujiu Yang

机构 * Tsinghua University(清华大学) ByteDance(字节跳动) Zhejiang University(浙江大学) Ping An Technology (Shenzhen) Co., Ltd(平安科技(深圳)有限公司)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025 Main Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16296 2025-10-07 cs.AI 83%

Cross-Modal Distillation For Widely Differing Modalities

Cairong Zhao, Yufeng Jin, Zifan Song, Haonan Chen, Duoqian Miao, Guosheng Hu

机构 * Department of Computer Science & Technology, Tongji University(计算机科学与技术系,同济大学) Alibaba Group(阿里巴巴集团) Oosto

专题命中 音频语音多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04136 2025-10-07 eess.AS cs.CV cs.SD 81%

MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition

Umberto Cappellazzo, Minsu Kim, Pingchuan Ma, Honglie Chen, Xubo Liu, Stavros Petridis, Maja Pantic

机构 * Imperial College London(伦敦帝国学院) Meta AI NatWest AI Research(NatWest人工智能研究)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07748 2025-10-07 eess.AS 79%

Leveraging Self-Supervised Audio-Visual Pretrained Models to Improve Vocoded Speech Intelligibility in Cochlear Implant Simulation

Richard Lee Lai, Jen-Cheng Hou, I-Chun Chern, Kuo-Hsuan Hung, Yi-Ting Chen, Mandar Gogate, Tughrul Arslan, Amir Hussain, Yu Tsao

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04738 2025-10-07 cs.SD cs.AI cs.CL cs.LG eess.AS 67%

Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba

Baher Mohammad, Magauiya Zhussip, Stamatios Lefkimmiatis

机构 * MTS AI(MTS人工智能公司) ITMO University(ITMO大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03639 2025-10-07 cs.CL cs.AI 62%

Towards Unsupervised Speech Recognition at the Syllable-Level

Liming Wang, Junrui Ni, Kai-Wei Chang, Saurabhchand Bhati, David Harwath, Mark Hasegawa-Johnson, James R. Glass

机构 * Massachusetts Institute of Technology(麻省理工学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04750 2025-10-07 cs.CL cs.SE 57%

A Low-Resource Speech-Driven NLP Pipeline for Sinhala Dyslexia Assistance

Peshala Perera, Deshan Sumanathilaka

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 11 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03336 2025-10-07 cs.SD cs.AI cs.LG 57%

Linguistic and Audio Embedding-Based Machine Learning for Alzheimer's Dementia and Mild Cognitive Impairment Detection: Insights from the PROCESS Challenge

Adharsha Sam Edwin Sam Devahi, Sohail Singh Sangha, Prachee Priyadarshinee, Jithin Thilakan, Ivan Fu Xing Tan, Christopher Johann Clarke, Sou Ka Lon, Balamurali B T, Yow Wei Quin, Chen Jer-Ming

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Hochschule für Musik Detmold(音乐学院Detmold)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2510.03727 2025-10-07 cs.AI cs.CL cs.CV cs.LG 89%

Bridging the Gap Between Multimodal Foundation Models and World Models

Xuehai He

机构 * Computer Science and Engineering University of California, Santa Cruz(计算机科学与工程大学加州大学圣克ruz分校)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06461 2025-10-07 cs.CV 79%

Interactive Test-Time Adaptation with Reliable Spatial-Temporal Voxels for Multi-Modal Segmentation

Haozhi Cao, Yuecong Xu, Pengyu Yin, Xingyu Ji, Shenghai Yuan, Jianfei Yang, Lihua Xie

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04819 2025-10-07 cs.CV cs.CL 62%

Visual Representations inside the Language Model

Benlin Liu, Amita Kamath, Madeleine Grunde-McLaughlin, Winson Han, Ranjay Krishna

机构 * University of Washington(华盛顿大学) University of California Los Angeles(加州大学洛杉矶分校) Allen Institute for AI(人工智能研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04753 2025-10-07 cs.CV 57%

Beyond Appearance: Transformer-based Person Identification from Conversational Dynamics

Masoumeh Chapariniya, Teodora Vukovic, Sarah Ebling, Volker Dellwo

机构 * Department of Computational Linguistics, University of Zurich, Zurich, Switzerland(计算语言学系,苏黎世大学,苏黎世,瑞士)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11336 2025-10-07 cs.CV 57%

UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks

Peiran Wu, Yunze Liu, Zhengdong Zhu, Enmin Zhou, Junxiao Shen

机构 * University of Bristol(布里斯托大学) Memories.ai Research(Memories.ai研究)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 7 篇

2510.03458 2025-10-07 cs.CL 83%

Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video

Mengyao Xu, Wenfei Zhou, Yauhen Babakhin, Gabriel Moreira, Ronay Ak, Radek Osmulski, Bo Liu, Even Oldridge, Benedikt Schifferer

机构 * NVIDIA

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11465 2025-10-07 cs.CL cs.LG 79%

CEMTM: Contextual Embedding-based Multimodal Topic Modeling

Amirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty, Giuseppe Carenini

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03309 2025-10-07 cs.LG q-bio.BM 67%

Thin Bridges for Drug Text Alignment: Lightweight Contrastive Learning for Target Specific Drug Retrieval

Mallikarjuna Tupakula

机构 * Rochester Institute of Technology(罗切斯特技术研究所)

专题命中 跨模态检索 :multimodal(abstract);multimodal foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏