arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-12 至 2025-11-12 共收录 52 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2504.14359 2025-11-12 cs.CV cs.AI cs.CL 82%

A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling

Kyle Buettner, Jacob T. Emmerson, Adriana Kovashka

机构 * Intelligent Systems Program(智能系统项目) Department of Computer Science(计算机科学系)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at IJCNLP-AACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07966 2025-11-12 cs.CV 79%

Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection

Shenao Zhao, Pengpeng Liang, Zhoufan Yang

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18711 2025-11-12 cs.CV cs.AI 62%

RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images

Ke Li, Di Wang, Ting Wang, Fuyu Dong, Yiming Zhang, Luyao Zhang, Xiangyu Wang, Shaofeng Li, Quan Wang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments This work is accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17352 2025-11-12 cs.CV cs.CL 62%

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Yihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng, Wei Wang, Kai-Wei Chang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 23 pages, 11 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2511.08031 2025-11-12 cs.CV cs.AI 86%

Multi-modal Deepfake Detection and Localization with FPN-Transformer

Chende Zheng, Ruiqi Suo, Zhoulin Ji, Jingyi Deng, Fangbin Yi, Chenhao Lin, Chao Shen

机构 * Xi’an Jiaotong University(西安交通大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13630 2025-11-12 cs.CV 85%

AVAR-Net: A Lightweight Audio-Visual Anomaly Recognition Framework with a Benchmark Dataset

Amjid Ali, Zulfiqar Ahmad Khan, Altaf Hussain, Muhammad Munsif, Adnan Hussain, Sung Wook Baik

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments I would like to request the withdrawal of my paper . The reason for this request is that I am currently working on additional experiments and analyses, which will lead to updates in the results section. Once these updates are complete, I will resubmit the revised version. Thank you for your understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10069 2025-11-12 cs.DC cs.LG 82%

ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism

Zedong Liu, Shenggan Cheng, Guangming Tan, Yang You, Dingwen Tao

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Electronic Science and Technology of China(电子科技大学) National University of Singapore(新加坡国立大学)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract)

Comments Accepted at NeurIPS 2025 Oral (Thirty-Ninth Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19281 2025-11-12 cs.RO 82%

Audio-Visual Traffic Light State Detection for Urban Robots

Sagar Gupta, Akansel Cosgun

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);multi-modal(abstract)

Comments Submitted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2024

Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13979 2025-11-12 cs.CL 79%

Mixed Signals: Understanding Model Disagreement in Multimodal Empathy Detection

Maya Srikanth, Run Chen, Julia Hirschberg

机构 * Columbia University(哥伦比亚大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments To appear in Findings of IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09205 2025-11-12 cs.MM cs.CL cs.IR cs.SD eess.AS 67%

Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model

Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper

机构 * University of Rochester(罗切斯特大学) Microsoft Research(微软研究院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.MM、eess.AS

Comments Accepted at EUSIPCO 2025 - 5 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 3 篇

2511.07988 2025-11-12 cs.AI 83%

The One Where They Brain-Tune for Social Cognition: Multi-Modal Brain-Tuning on Friends

Nico Policzer, Cameron Braunstein, Mariya Toneva

机构 * Saarland University(萨尔兰大学) MPI for Software Systems(软件系统Max Planck研究所) University of British Columbia(不列颠哥伦比亚大学)

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments 20 pages, 7 figures. Appearing at the NeurIPS 2025 Workshop on Interpreting Cognition in Deep Learning Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06958 2025-11-12 cs.CV 70%

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Xinhao Li, Ziang Yan, Desen Meng, Lu Dong, Xiangyu Zeng, Yinan He, Yali Wang, Yu Qiao, Yi Wang, Limin Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Nanjing University(南京大学) Zhejiang University(浙江大学) University of Science and Technology of China(中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08535 2025-11-12 cs.CV cs.AI 62%

Large Sign Language Models: Toward 3D American Sign Language Translation

Sen Zhang, Xiaoxiao He, Di Liu, Zhaoyang Xia, Mingyu Zhao, Chaowei Tan, Vivian Li, Bo Liu, Dimitris N. Metaxas, Mubbasir Kapadia

机构 * Rutgers University(罗格斯大学) Meta Reality Labs(Meta现实实验室) Qualcomm(高通公司) Walmart Global Tech(沃尔玛全球技术) Roblox PRISMS

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2511.08369 2025-11-12 cs.CV cs.AI 62%

Text-based Aerial-Ground Person Retrieval

Xinyu Zhou, Yu Wu, Jiayao Ma, Wenhao Wang, Min Cao, Mang Ye

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14772 2025-11-12 cs.HC 50%

UMind: A Unified Multitask Network for Zero-Shot M/EEG Visual Decoding

Chengjian Xu, Yonghao Song, Zelin Liao, Haochuan Zhang, Qiong Wang, Qingqing Zheng

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2510.14631 2025-11-12 cs.DB 78%

Towards a Multimodal Stream Processing System

Uélison Jean Lopes dos Santos, Alessandro Ferri, Szilard Nistor, Riccardo Tommasini, Carsten Binnig, Manisha Luthra

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07934 2025-11-12 cs.CV 74%

Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers

Sida Huang, Siqi Huang, Ping Luo, Hongyuan Zhang

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07816 2025-11-12 cs.CV 74%

Cancer-Net PCa-MultiSeg: Multimodal Enhancement of Prostate Cancer Lesion Segmentation Using Synthetic Correlated Diffusion Imaging

Jarett Dewbury, Chi-en Amy Tai, Alexander Wong

机构 * Systems Design Engineering University of Waterloo(水力工程系统设计系大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted at ML4H 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08536 2025-11-12 cs.CV 57%

3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation

Yunhong He, Zhengqing Yuan, Zhengzhong Tu, Yanfang Ye, Lichao Sun

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07877 2025-11-12 cs.CV 57%

Visual Bridge: Universal Visual Perception Representations Generating

Yilin Gao, Shuguang Dou, Junzhou Li, Zhiheng Yu, Yin Li, Dongsheng Jiang, Shugong Xu

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07744 2025-11-12 cs.CV 57%

VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics

Daniel Cher, Brian Wei, Srikumar Sastry, Nathan Jacobs

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07690 2025-11-12 cs.AI 57%

Towards AI-Assisted Generation of Military Training Scenarios

Soham Hans, Volkan Ustun, Benjamin Nye, James Sterrett, Matthew Green

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03668 2025-11-12 cs.DC cs.LG 50%

Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge

Fernando Koch, Aladin Djuhera, Alecio Binotto

机构 * Florida Atlantic University, USA(佛罗里达大学) Technical University Munich, Germany(慕尼黑技术大学) Carl Zeiss AG, Germany(蔡司股份公司)

专题命中 多模态生成 :multi-modal(abstract)

Comments 26 pages, 3 figures, 4 tables, 52 references

Journal ref Computer Networks and Communications, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 11 篇

2511.08263 2025-11-12 cs.CV cs.AI 84%

ImagebindDC: Compressing Multi-modal Data with Imagebind-based Condensation

Yue Min, Shaobo Wang, Jiaze Li, Tianle Niu, Junxin Fan, Yongliang Miao, Lijin Yang, Linfeng Zhang

机构 * EPIC Lab, SJTU(上海交通大学EPIC实验室) Bosch Corporate Research Asia Pacific(博世亚太公司研究部) HKUST(香港科技大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments AAAI 2026, 18 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07812 2025-11-12 cs.CV 83%

Revisiting MLLM Based Image Quality Assessment: Errors and Remedy

Zhenchen Tang, Songlin Yang, Bo Peng, Zichuan Wang, Jing Dong

专题命中 多模态评测 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13231 2025-11-12 cs.CV cs.AI 81%

WildFireCan-MMD: A Multimodal Dataset for Classification of User-Generated Content During Wildfires in Canada

Braeden Sherritt, Isar Nejadgholi, Efstratios Aivaliotis, Khaled Mslmani, Marzieh Amini

机构 * Carleton University(卡尔顿大学) National Research Council Canada(加拿大国家研究委员会)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22262 2025-11-12 cs.CV 79%

UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data

Yujian Yuan, Changjie Wu, Xinyuan Chang, Sijin Wang, Hang Zhang, Shiyi Liang, Shuang Zeng, Mu Xu, Ning Guo

机构 * Alibaba Group(阿里巴巴集团)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments AAAI2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13861 2025-11-12 cs.HC cs.CL cs.MA 79%

3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark

Ivan Sviridov, Amina Miftakhova, Artemiy Tereshchenko, Galina Zubkova, Pavel Blinov, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室) HSE University(俄罗斯高等经济大学) ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与系统问题研究所可信人工智能研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments EMNLP 25 (main)

Journal ref https://aclanthology.org/2025.emnlp-main.1353/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07777 2025-11-12 eess.SP 78%

A Causal-Guided Multimodal Large Language Model for Generalized Power System Time-Series Data Analytics

Zhenghao Zhou, Yiyan Li, Xinjie Yu, Runlong Liu, Zelin Guo, Zheng Yan, Mo-Yuen Chow, Yuqi Yang, Yang Xu

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08133 2025-11-12 cs.CV cs.AI 73%

OTSNet: A Neurocognitive-Inspired Observation-Thinking-Spelling Pipeline for Scene Text Recognition

Lixu Sun, Nurmemet Yolwas, Wushour Silamu

机构 * School of Computer Science and Technology(计算机科学与技术学院) School of Computer Science(计算机科学学院) Xinjiang University(新疆大学)

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏