arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-29 至 2025-10-29 共收录 60 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2510.24331 2025-10-29 cs.LG cs.CV 79%

What do vision-language models see in the context? Investigating multimodal in-context learning

Gabriel O. dos Santos, Esther Colombini, Sandra Avila

机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP)(计算机学院,Campinas州立大学(UNICAMP))

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24446 2025-10-29 cs.CL cs.CV 62%

SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space

Viktoriia Zinkovich, Anton Antonov, Andrei Spiridonov, Denis Shepelev, Andrey Moskalenko, Daria Pugacheva, Elena Tutubalina, Andrey Kuznetsov, Vlad Shakhuro

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24321 2025-10-29 cs.CV cs.AI 62%

Few-Shot Remote Sensing Image Scene Classification with CLIP and Prompt Learning

Ivica Dimitrovski, Vlatko Spasev, Ivan Kitanovski

机构 * Faculty of Computer Science and Engineering(计算机科学与工程学院) University Ss Cyril and Methodius(西里尔与美多西乌斯大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05229 2025-10-29 cs.CV cs.MM 62%

Does CLIP perceive art the same way we do?

Andrea Asperti, Leonardo Dessì, Maria Chiara Tonetti, Nico Wu

机构 * Dept. of Informatics (DISI) University of Bologna(信息学院(DISI)博洛尼亚大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Journal ref Proceedings of IEEE International Conference on Content-Based Multimedia Indexing (IEEE CBMI 2025), Dublin, Ireland, 22-24 October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24650 2025-10-29 cs.AI 57%

Advancing site-specific disease and pest management in precision agriculture: From reasoning-driven foundation models to adaptive, feedback-based learning

Nitin Rai, Daeun, Choi, Nathan S. Boyd, Arnold W. Schumann

机构 * Department of Horticultural Sciences(园艺科学系) Gulf Coast Research and Education Center(墨西哥湾沿岸研究与教育中心) University of Florida(佛罗里达大学) Department of Agricultural and Biological Engineering(农业与生物工程系) Department of Soil, Water, and Ecosystem Sciences(土壤、水与生态系统科学系) Citrus Research and Education Center(柑橘研究与教育中心)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.AI

Comments 26 pages, 8 figures, and 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07046 2025-10-29 cs.CV 57%

RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning

Yunchuan Ma, Laiyun Qing, Guorong Li, Yuankai Qi, Amin Beheshti, Quan Z. Sheng, Qingming Huang

机构 * University of Chinese Academy of Science, Beijing,100190, China(中国科学院大学) Macquarie University(麦考瑞大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Published in Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23894 2025-10-29 cs.CV 57%

Improving Visual Discriminability of CLIP for Training-Free Open-Vocabulary Semantic Segmentation

Jinxin Zhou, Jiachen Jiang, Zhihui Zhu

机构 * Department of Computer Science(计算机科学系)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 23 pages, 10 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15256 2025-10-29 cs.CV 57%

Normal and Abnormal Pathology Knowledge-Augmented Vision-Language Model for Anomaly Detection in Pathology Images

Jinsol Song, Jiamu Wang, Anh Tien Nguyen, Keunho Byeon, Sangjeong Ahn, Sung Hak Lee, Jin Tae Kwak

机构 * Korea University(韩国大学) The Catholic University of Korea(韩国天主大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Code is available at: https://github.com/QuIIL/ICCV2025_Ano-NAViLa

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 8 篇

2505.21724 2025-10-29 cs.CV cs.AI cs.HC 86%

OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions

Cheng Luo, Jianghui Wang, Bing Li, Siyang Song, Bernard Ghanem

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);audio-visual(abstract);分类 cs.CV、cs.AI

Comments 25 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24024 2025-10-29 eess.AS cs.CV eess.IV 81%

Listening without Looking: Modality Bias in Audio-Visual Captioning

Yuchi Ishikawa, Toranosuke Manabe, Tatsuya Komatsu, Yoshimitsu Aoki

机构 * LY Corporation(LY公司) Keio University(庆应大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24247 2025-10-29 cs.CL 79%

Abjad AI at NADI 2025: CATT-Whisper: Multimodal Diacritic Restoration Using Text and Speech Representations

Ahmad Ghannam, Naif Alharthi, Faris Alasmary, Kholood Al Tabash, Shouq Sadah, Lahouari Ghouti

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24103 2025-10-29 cs.SD cs.AI cs.MM eess.AS 75%

Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation

Kang Zhang, Trung X. Pham, Suyeon Lee, Axi Niu, Arda Senocak, Joon Son Chung

机构 * KAIST, South Korea(韩国釜山国立大学) NWPU, China(中国西北工业大学) UNIST, South Korea(韩国全南国立大学)

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.AI、cs.MM、eess.AS

Comments accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24628 2025-10-29 cs.CL 57%

"Mm, Wat?" Detecting Other-initiated Repair Requests in Dialogue

Anh Ngo, Nicolas Rollet, Catherine Pelachaud, Chloe Clavel

机构 * ALMAnaCH, INRIA Paris(ALMAnaCH,INRIA巴黎) Télécom Paris, SES, Institut Polytechnique de Paris, I3-CNRS(Télécom巴黎,SES,巴黎高等理工学院,I3-CNRS) Télécom Paris, LTCI, Institut Polytechnique de Paris(Télécom巴黎,LTCI,巴黎高等理工学院) CNRS, ISIR, Sorbonne University(CNRS,ISIR,索邦大学) ISIR, Sorbonne University(ISIR,索邦大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18195 2025-10-29 eess.AS cs.LG cs.SD 57%

Acoustic and Machine Learning Methods for Speech-Based Suicide Risk Assessment: A Systematic Review

Ambre Marie, Marine Garnier, Thomas Bertin, Laura Machart, Guillaume Dardenne, Gwenolé Quellec, Sofian Berrouiguet

机构 * University of Western Brittany(西部布列塔尼大学) Psychiatry Department, CHU Brest(布列塔尼大学精神科部门)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Preprint version of a manuscript submitted to the Journal of Affective Disorders

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23887 2025-10-29 cs.HC 50%

MORA: AI-Mediated Story-Based practice for Speech Sound Disorder from Clinic to Home

Sumin Hong, Xavier Briggs, Qingxiao Zheng, Yao Du, Jinjun Xiong, Toby Jia-jun Li

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24180 2025-10-29 cs.LG 50%

V-SAT: Video Subtitle Annotation Tool

Arpita Kundu, Joyita Chakraborty, Anindita Desarkar, Aritra Sen, Srushti Anil Patil, Vishwanathan Raman

机构 * LTIMindTree

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 9 篇

2510.15870 2025-10-29 cs.CV cs.AI cs.CL 85%

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

Hanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang, Ligeng Zhu, Yuanhang Su, Sean Lin, An-Chieh Cheng, Zhen Wan, Jinchuan Tian, Yuming Lou, Dong Yang, Zhijian Liu, Yukang Chen, Ambrish Dantrey, Ehsan Jahangiri, Sreyan Ghosh, Daguang Xu, Ehsan Hosseini-Asl, Danial Mohseni Taheri, Vidya Murali, Sifei Liu, Yao Lu, Oluwatobi Olabiyi, Yu-Chiang Frank Wang, Rafael Valle, Bryan Catanzaro, Andrew Tao, Song Han, Jan Kautz, Hongxu Yin, Pavlo Molchanov

机构 * NVIDIA

专题命中 视频多模态 :omni-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical Report. Code: https://github.com/NVlabs/OmniVinci

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17637 2025-10-29 cs.LG 82%

Causal Spatio-Temporal Prediction: An Effective and Efficient Multi-Modal Approach

Yuting Huang, Ziquan Fang, Zhihao Zeng, Lu Chen, Yunjun Gao

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23934 2025-10-29 cs.CY cs.AI cs.ET 79%

MFiSP: A Multimodal Fire Spread Prediction Framework

Alec Sathiyamoorthy, Wenhao Zhou, Xiangmin Zhou, Xiaodong Li, Iqbal Gondal

机构 * School of Computing Technologies, RMIT University(计算技术学院,皇家墨尔本理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03727 2025-10-29 eess.AS cs.LG 79%

Detecting Neurocognitive Disorders through Analyses of Topic Evolution and Cross-modal Consistency in Visual-Stimulated Narratives

Jinchao Li, Yuejiao Wang, Junan Li, Jiawen Kang, Bo Zheng, Ka Ho Wong, Brian Mak, Helene H. Fung, Jean Woo, Man-Wai Mak, Timothy Kwok, Vincent Mok, Xianmin Gong, Xixin Wu, Xunying Liu, Patrick C. M. Wong, Helen Meng

机构 * The Chinese University of Hong Kong(香港中文大学) The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 eess.AS

Comments 16 pages, 5 figures, accepted by "IEEE Journal of Selected Topics in Signal Processing"

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24238 2025-10-29 cond-mat.mtrl-sci 78%

Unlocking Dynamic Luminescent Mapping of pH with Sustainable Lignin-Derived Carbon Dots with Multimodal Readout Capacity

Maja Szymczak, Jan Hočevar, Jernej Iskra, Darja Lisjak, Jelena Papan Djaniš, Lukasz Marciniak, Karolina Elzbieciak-Piecka

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23630 2025-10-29 cs.LG cs.AI cs.CL 62%

NUM2EVENT: Interpretable Event Reasoning from Numerical time-series

Ninghui Feng, Yiyan Qi

机构 * International Digital Economy Academy (IDEA)(国际数字经济学院(IDEA)) University of Nottingham Ningbo(诺丁汉大学宁波校区)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22092 2025-10-29 cs.AI 57%

VIRAL: Vision-grounded Integration for Reward design And Learning

Valentin Cuzin-Rambaud, Emilien Komlenovic, Alexandre Faure, Bruno Yun

机构 * Université Claude Bernard Lyon 1(克莱尔伯恩大学里昂1分校) CNRS(国家科学研究中心) Ecole Centrale de Lyon(里昂中央理工学院) INSA Lyon(里昂工业高等学院) Université Lumière Lyon 2(里昂2大学卢米埃尔分校) LIRIS(图像研究所)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24055 2025-10-29 cs.RO cs.LG 50%

Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation

Xiucheng Zhang, Yang Jiang, Hongwei Qing, Jiashuo Bai

机构 * Xiucheng Zhang(未知) Yang Jiang(未知) Jiashuo Bai(未知)

专题命中 视频多模态 :multimodal(abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05868 2025-10-29 cs.SI 50%

Detecting Coordinated Behaviour on Video-First Platforms: The Challenge of Multimodality and Complex Similarity on TikTok

Inga K. Wohlert, Davide Vega, Matteo Magnani, Alexandra Segerberg

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 1 篇

2510.23617 2025-10-29 cs.LG cs.AI cs.CL 84%

An Enhanced Dual Transformer Contrastive Network for Multimodal Sentiment Analysis

Phuong Q. Dao, Mark Roantree, Vuong M. Ngo

机构 * Ho Chi Minh City Open University, Ho Chi Minh, Vietnam Insight Centre for Data Analytics \& School of Computing, Dublin City University, Dublin, Ireland

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments The paper has been accepted for presentation at the MEDES 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 3 篇

2510.24514 2025-10-29 cs.CV cs.CL 81%

Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs

Huanyu Zhang, Wenshan Wu, Chengzu Li, Ning Shang, Yan Xia, Yangyu Huang, Yifan Zhang, Li Dong, Zhang Zhang, Liang Wang, Tieniu Tan, Furu Wei

机构 * MSR(微软研究院) UCAS(中国科学院自动化研究所) CASIA(中国科学院自动化研究所) Cambridge(剑桥大学) NJU(南京大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16691 2025-10-29 cs.CV 57%

InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention

Qiang Xiang, Shuang Sun, Binglei Li, Dejia Song, Huaxia Li, Nemo Chen, Xu Tang, Yao Hu, Junping Zhang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24083 2025-10-29 math.OC 50%

A Novel Virus Diffusion Optimization (VDO) Algorithm for Global Optimization

Zhaoqi Sun, Qingsong Wang

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 12 篇

2506.17939 2025-10-29 cs.CV cs.AI 84%

GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning

Bo Liu, Xiangyu Zhao, Along He, Yidi Chen, Huazhu Fu, Xiao-Ming Wu

机构 * The Hong Kong Polytechnic University(香港理工大学) Shenzhen University(深圳大学) West China Hospital of Sichuan University(四川大学华西医院) IHPC, Agency for Science, Technology and Research(科技研究局IHPC)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ACM MM 2025 (also known as GEMeX-ThinkVG)

详情

展开后加载摘要…

URL PDF HTML 收藏