arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-30 至 2025-10-30 共收录 42 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2506.21710 2025-10-30 cs.CV 86%

FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering

Liangyu Zhong, Fabio Rosenthal, Joachim Sicking, Fabian Hüger, Thorsten Bagdonat, Hanno Gottschalk, Leo Schwinn

机构 * Technical University of Berlin(柏林技术大学) Technical University of Munich(慕尼黑技术大学) CARIAD SE Volkswagen AG(大众集团)

专题命中 图文多模态 :MLLM(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 - main track. Project page: https://focus-mllm-vqa.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04508 2025-10-30 cs.CL 85%

Adapter-state Sharing CLIP for Parameter-efficient Multimodal Sarcasm Detection

Soumyadeep Jana, Sahil Danayak, Sanasam Ranbir Singh

机构 * Department of Computer Science(计算机科学系) Engineering, Indian Institute of Technology Guwahati, India(工程系、印度理工学院瓜哇提学院、印度)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19311 2025-10-30 cs.CV cs.AI 84%

DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment

Weizhi Chen, Yupeng Deng, Jin Wei, Jingbo Chen, Jiansheng Chen, Yuman Feng, Zhihao Xi, Diyou Liu, Kai Li, Yu Meng

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院 aerospace information research institute) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) School of Information Network Security, People’s Public Security University of China(中国人民公安大学信息网络安全学院)

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25303 2025-10-30 cs.CL 83%

Teaching Sarcasm: Few-Shot Multimodal Sarcasm Detection via Distillation to a Parameter-Efficient Student

Soumyadeep Jana, Sanasam Ranbir Singh

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Indian Institute of Technology Guwahati(印度理工学院古瓦哈蒂)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25070 2025-10-30 cs.CV 70%

Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments

Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Preprint under review at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25179 2025-10-30 cs.AI 57%

Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University, Australia(计算机学院,麦考瑞大学,澳大利亚)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25175 2025-10-30 cs.CV 57%

Test-Time Adaptive Object Detection with Foundation Model

Yingjie Gao, Yanan Zhang, Zhi Cai, Di Huang

机构 * State Key Laboratory of Complex and Critical Software Environment, Beihang University(复杂与关键软件环境国家重点实验室,北京航空航天大学) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25051 2025-10-30 cs.CV cs.LG 57%

Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models

Shunjie-Fabian Zheng, Hyeonjun Lee, Thijs Kooi, Ali Diba

机构 * Department of Medicine I, LMU University Hospital, LMU Munich, Germany(慕尼黑大学医学院第一医学部,LMU慕尼黑大学医院) Lunit Inc.(Lunit公司)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to Computer Vision for Automated Medical Diagnosis (CVAMD) Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2501.14755 2025-10-30 cs.DC cs.AI 57%

Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models

Daoyuan Chen, Yilun Huang, Xuchen Pan, Nana Jiang, Haibin Wang, Yilei Zhang, Ce Ge, Yushuo Chen, Wenhao Zhang, Zhijian Ma, Jun Huang, Wei Lin, Yaliang Li, Bolin Ding, Jingren Zhou

机构 * Alibaba Group(阿里巴巴集团)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 (Spotlight). 43 pages, 16 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25199 2025-10-30 cs.CV 57%

AI-Powered Early Detection of Critical Diseases using Image Processing and Audio Analysis

Manisha More, Kavya Bhand, Kaustubh Mukdam, Kavya Sharma, Manas Kawtikwar, Hridayansh Kaware, Prajwal Kavhar

机构 * Dept. of Computer Engineering(计算机工程系) Vishwakarma Institute of Technology(维斯瓦卡arma技术学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19902 2025-10-30 cs.CL 57%

WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction

Binbin Zhang, Chengdong Liang, Shuai Wang, Xuelong Geng, Zhao Guo, Haoyu Li, Hao Yin, Xipeng Yang, Pengshen Zhang, Changwei Ma, Lei Xie

机构 * Northwestern Polytechnical University(西北工业大学) Nanjing University(南京大学) Shanghai Jiao Tong University(上海交通大学) GuaSemi Speech A Team(GuaSemi语音团队) WeNet Community(WeNet社区)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2510.25091 2025-10-30 cs.AI 83%

H3M-SSMoEs: Hypergraph-based Multimodal Learning with LLM Reasoning and Style-Structured Mixture of Experts

Peilin Tan, Liang Xie, Churan Zhi, Dian Tu, Chuanqi Shi

机构 * University of California, San Diego(加州大学圣迭戈分校) Wuhan University of Technology(武汉科技大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17897 2025-10-30 q-bio.NC cs.CV cs.LG 79%

Multimodal Recurrent Ensembles for Predicting Brain Responses to Naturalistic Movies (Algonauts 2025)

Semih Eren, Deniz Kucukahmetler, Nico Scherf

机构 * Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所) TU Dresden(德累斯顿技术大学) School for Embedded and Composite AI (SECAI)(嵌入式与复合人工智能学院) Center for Scalable Data Analytics & AI (ScaDS.AI)(可扩展数据与人工智能中心)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 2 figures, 1 table. Invited report, CCN 2025 Algonauts Project session (3rd-place team). Code: https://github.com/erensemih/Algonauts2025_ModalityRNN v3: Added equal contribution footnote to author list. Corrected reference list

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25332 2025-10-30 cs.CV 79%

StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA

Yuhang Hu, Zhenyu Yang, Shihan Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Changsheng Xu

机构 * Henan Institute of Advanced Technology, Zhengzhou University(河南高级技术研究所,郑州大学) Institute of Automation, CAS(自动化研究所,中国科学院) UCAS(中国科学院大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24720 2025-10-30 cs.HC cs.AI cs.CV 62%

Modelling the Interplay of Eye-Tracking Temporal Dynamics and Personality for Emotion Detection in Face-to-Face Settings

Meisam J. Seikavandi, Jostein Fimland, Fabricio Batista Narcizo, Maria Barrett, Ted Vucurevich, Jesper Bünsow Boldt, Andrew Burke Dittberner, Paolo Burelli

机构 * brAIn Lab, IT University of Copenhagen(brAIn实验室,丹麦技术大学) GN Advanced Science(GN先进科学) IT University of Copenhagen(丹麦技术大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2509.10266 2025-10-30 cs.CV cs.AI 76%

SignMouth: Leveraging Mouthing Cues for Sign Language Translation by Multimodal Contrastive Fusion

Wenfang Wu, Tingting Yuan, Yupeng Li, Daling Wang, Xiaoming Fu

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25718 2025-10-30 cs.IR cs.DL 50%

Retrieval-Augmented Search for Large-Scale Map Collections with ColPali

Jamie Mahowald, Benjamin Charles Germain Lee

专题命中 跨模态检索 :multimodal(abstract)

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 4 篇

2510.24820 2025-10-30 cs.CV cs.AI 84%

SafeEditor: Unified MLLM for Efficient Post-hoc T2I Safety Editing

Ruiyang Zhang, Jiahao Luo, Xiaoru Feng, Qiufan Pang, Yaodong Yang, Juntao Dai

机构 * PKU Alignment Team, Peking University(北京大学对齐团队) LLM Safety Centre, Beijing Academy of Artificial Intelligence(北京人工智能研究院大语言模型安全中心)

专题命中 多模态生成 :MLLM(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22439 2025-10-30 cs.SD cs.AI 74%

PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching

Ali Vosoughi, Yongyi Zang, Qihui Yang, Nathan Paek, Randal Leistikow, Chenliang Xu

机构 * Smule Labs(Smule实验室) University of California, San Diego(加州大学圣地亚哥分校) University of Rochester(罗切斯特大学) Stanford University(斯坦福大学)

专题命中 多模态生成 :multimodal(title);分类 cs.AI

Comments 9 pages, 2 figures, 4 tables; v2: corrected spelling of a co-author name; no content changes

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25163 2025-10-30 cs.CV 57%

Target-Guided Bayesian Flow Networks for Quantitatively Constrained CAD Generation

Wenhao Zheng, Chenwei Sun, Wenbo Zhang, Jiancheng Lv, Xianggen Liu

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, Chengdu, China(教育部机器学习与工业智能工程研究中心)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (2025) 3330-3339

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20071 2025-10-30 cs.HC 50%

Towards Human-AI Synergy in UI Design: Supporting Iterative Generation with LLMs

Mingyue Yuan, Jieshan Chen, Yongquan Hu, Sidong Feng, Mulong Xie, Gelareh Mohammadi, Zhenchang Xing, Aaron Quigley

专题命中 多模态生成 :multi-modal(abstract)

Comments ACM Transactions on Computer-Human Interaction

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 8 篇

2510.24980 2025-10-30 cs.CV cs.AI 84%

FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning

Reza Saadati Fard, Emmanuel Agu, Palawat Busaranuvong, Deepak Kumar, Shefalika Gautam, Bengisu Tulu, Diane Strong, Lorraine Loretz

机构 * Worcester Polytechnic Institute(沃斯特理工学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04458 2025-10-30 cs.CL 83%

Think Twice Before You Judge: Mixture of Dual Reasoning Experts for Multimodal Sarcasm Detection

Soumyadeep Jana, Abhrajyoti Kundu, Sanasam Ranbir Singh

机构 * Indian Institute of Technology Guwahati(印度理工学院古瓦哈提)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25120 2025-10-30 cs.SI 82%

MMM-Fact: A Multimodal, Multi-Domain Fact-Checking Dataset with Multi-Level Retrieval Difficulty

Wenyan Xu, Dawei Xiang, Tianqi Ding, Weihai Lu

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

Comments Dataset link: https://huggingface.co/datasets/Wenyan0110/MMM-Fact

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19028 2025-10-30 cs.CV cs.AI 81%

InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts

Tianchi Xie, Minzhi Lin, Mengchen Liu, Yilin Ye, Changjian Chen, Shixia Liu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25211 2025-10-30 cs.RO 71%

RoadSens-4M: A Multimodal Smartphone & Camera Dataset for Holistic Road-way Analysis

Amith Khandakar, David Michelson, Shaikh Golam Rabbani, Fariya Bintay Shafi, Md. Faysal Ahamed, Khondokar Radwanur Rahman, Md Abidur Rahman, Md. Fahmidun Nabi, Mohamed Arselene Ayari, Khaled Khan, Ponnuthurai Nagaratnam Suganthan

专题命中 多模态评测 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25053 2025-10-30 cs.RO cs.AI cs.LG q-bio.NC 57%

Scalable predictive processing framework for multitask caregiving robots

Hayato Idei, Tamon Miyake, Tetsuya Ogata, Yuichi Yamashita

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14926 2025-10-30 cs.AI 57%

TraveLLM: Could you plan my new public transit route in face of a network disruption?

Bowen Fang, Zixiao Yang, Xuan Di

机构 * Department of Industrial Engineering and Operations Research, Columbia University(工业工程与运营管理系,哥伦比亚大学) Department of Civil Engineering and Engineering Mechanics, Columbia University(土木工程与工程力学系,哥伦比亚大学) Data Science Institute, Columbia University(数据科学研究院,哥伦比亚大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments Accepted to ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25001 2025-10-30 stat.CO cs.LG 50%

Bayesian Neural Networks vs. Mixture Density Networks: Theoretical and Empirical Insights for Uncertainty-Aware Nonlinear Modeling

Riddhi Pratim Ghosh, Ian Barnett

机构 * Department of Mathematics and Statistics, Bowling Green State University(数学与统计学系,布恩维尔格林州立大学) Department of Biostatistics, University of Pennsylvania(生物统计学系,宾夕法尼亚大学)

专题命中 多模态评测 :multimodal(abstract)

Comments 20 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 3 篇

2510.25092 2025-10-30 cs.MA 82%

SeeingEye: Agentic Information Flow Unlocks Multimodal Reasoning In Text-only LLMs

Weijia Zhang, Zijia Liu, Haoru Li, Haoqi Chen, Jiaxuan You

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏