arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-13 至 2025-10-13 共收录 50 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 13 篇

2510.09302 2025-10-13 cs.CV cs.AI cs.CL 67%

CapGeo: A Caption-Assisted Approach to Geometric Reasoning

Yuying Li, Siyi Qian, Hao Liang, Leqi Zheng, Ruichuan An, Yongzhen Guo, Wentao Zhang

机构 * THU(清华大学) PKU(北京大学) Ant Group(蚂蚁集团)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments preprint, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08936 2025-10-13 cs.CV cs.AI 62%

RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos

Zixi Yang, Jiapeng Li, Muxi Diao, Yinuo Jing, Kongming Liang

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09361 2025-10-13 cs.CV 57%

BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception

Junyan Ye, Dongzhi Jiang, Jun He, Baichuan Zhou, Zilong Huang, Zhiyuan Yan, Hongsheng Li, Conghui He, Weijia Li

机构 * Sun Yat-sen University(中山大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) CUHK MMLab(香港中文大学多模态实验室) Peking University(北京大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted to 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12882 2025-10-13 cs.CL 57%

ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videos

Patrick Giedemann, Pius von Däniken, Jan Deriu, Alvaro Rodrigo, Anselmo Peñas, Mark Cieliebak

机构 * Zurich University of Applied Sciences(苏黎世应用科学大学) UNED NLP & IR Group(UNED自然语言处理与信息检索组)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00047 2025-10-13 cs.LG cs.AI 57%

TriP-LLM: A Tri-Branch Patch-wise Large Language Model Framework for Time-Series Anomaly Detection

Yuan-Cheng Yu, Yen-Chieh Ouyang, Chun-An Lin

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments Accepted version of the paper published in IEEE Access (2025). Licensed under a Creative Commons Attribution 4.0 License (CC BY 4.0). Published version available at IEEE Xplore

Journal ref IEEE Access, vol. 13, pp. 168643-168653, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 2 篇

2506.19433 2025-10-13 cs.CV cs.AI cs.CL 75%

Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System

Lixuan He, Haoyu Dong, Zhenxing Chen, Yangcheng Yu, Jie Feng, Yong Li

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The paper is currently under investigation regarding concerns of potential academic misconduct. While the investigation is ongoing, the authors have voluntarily requested to withdraw the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09507 2025-10-13 cs.CV cs.RO 70%

PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs

Zixin Zhang, Kanghao Chen, Xingwang Lin, Lutao Jiang, Xu Zheng, Yuanhuiyi Lyu, Litao Guo, Yinchuan Li, Ying-Cong Chen

机构 * HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) Beihang University(北航大学) Knowin

专题命中 多模态Agent :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 10 篇

2510.08606 2025-10-13 cs.CL cs.AI 88%

Centering Emotion Hotspots: Multimodal Local-Global Fusion and Cross-Modal Alignment for Emotion Recognition in Conversations

Yu Liu, Hanlei Shi, Haoxun Li, Yuqing Sun, Yuxuan Ding, Linlin Gong, Leyuan Qu, Taihao Li

机构 * Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(杭州高等研究 institute,中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CL、cs.AI

Comments Under review for ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21447 2025-10-13 cs.CV cs.AI 84%

Multimodal Language Models See Better When They Look Shallower

Haoran Chen, Junyan Lin, Xinghao Chen, Yue Fan, Jianfeng Dong, Xin Jin, Hui Su, Jinlan Fu, Xiaoyu Shen

机构 * Zhejiang Gongshang University(浙江工商大学) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波) Meituan Inc.(美团公司) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures, accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08589 2025-10-13 cs.CV cs.AI 84%

Beyond CNNs: Efficient Fine-Tuning of Multi-Modal LLMs for Object Detection on Low-Data Regimes

Nirmal Elamon, Rouzbeh Davoudi

机构 * Artificial Creative intelligence (ACI)(人工创意智能(ACI)) Expedia Group(Expedia集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.06145 2025-10-13 cs.CV cs.AI cs.HC cs.LG stat.ML 81%

Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training

Mahdi Abavisani, Hamid Reza Vaezi Joze, Vishal M. Patel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 1165-1174

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.06498 2025-10-13 cs.LG cs.AI cs.CV stat.ML 81%

Deep Multimodal Subspace Clustering Networks

Mahdi Abavisani, Vishal M. Patel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref IEEE Journal of Selected Topics in Signal Processing, vol. 12, no. 6, pp. 1601-1614, Dec. 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07979 2025-10-13 cs.CV 79%

Visual Representation Alignment for Multimodal Large Language Models

Heeji Yoon, Jaewoo Jung, Junwan Kim, Hyungyu Choi, Heeseong Shin, Sangbeom Lim, Honggyu An, Chaehyun Kim, Jisang Han, Donghyun Kim, Chanho Eom, Sunghwan Hong, Seungryong Kim

机构 * KAIST AI(韩国科学技术院人工智能研究所) New York University(纽约大学) Chung-Ang University(Chung-Ang 大学) Korea University(韩国大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://cvlab-kaist.github.io/VIRAL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02912 2025-10-13 cs.CV 70%

Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention

Xin Zou, Di Lu, Yizhou Wang, Yibo Yan, Yuanhuiyi Lyu, Xu Zheng, Linfeng Zhang, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08802 2025-10-13 cs.LG 67%

Edu-EmotionNet: Cross-Modality Attention Alignment with Temporal Feedback Loops

S M Rafiuddin

机构 * Department of Computer Science Oklahoma State University Stillwater, Oklahoma, USA(计算机科学系 奥克拉荷马州立大学 斯蒂尔沃特 奥克拉荷马州 美国)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract)

Comments 6 Pages, 6 Figures, 3 Tables, Accepted as a Regular Research paper at ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09435 2025-10-13 cs.LG cs.IR 50%

Cross-attention Secretly Performs Orthogonal Alignment in Recommendation Models

Hyunin Lee, Yong Zhang, Hoang Vu Nguyen, Xiaoyi Liu, Namyong Park, Christopher Jung, Rong Jin, Yang Wang, Zhigang Wang, Somayeh Sojoudi, Xue Feng

机构 * Meta

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25641 2025-10-13 physics.med-ph 50%

Semi-Supervised Radiomics for Glioblastoma IDH Mutation: Limited Labels, Data Sensitivity, and SHAP Interpretation

Amir Hossein Pouria, Shahram Taeb, Somayeh Sadat Mehrnia, Sajad Jabarzadeh Ghandilu, Mehrdad Oveisi, Arman Rahmim, Mohammad R. Salmanpour

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 13 Pages and 4 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 3 篇

2510.09513 2025-10-13 stat.ML cs.LG 78%

Interpretable Generative and Discriminative Learning for Multimodal and Incomplete Clinical Data

Albert Belenguer-Llorens, Carlos Sevilla-Salcedo, Janaina Mourao-Miranda, Vanessa Gómez-Verdejo

机构 * Department of Health Sciences and Technology (DHEST), ETH Zurich(健康科学与技术系,苏黎世联邦理工学院) Department of Signal Theory and Communications, Universidad Carlos III de Madrid(信号理论与通信系,卡洛斯三世大学) UCL Hawkes Institute, Department of Computer Science, University College London(UCL Hawkes研究所,伦敦大学学院计算机科学系) Instituto de Investigación Sanitaria Gregorio Marañón (IiSGM), Madrid(Gregorio Marañón卫生研究所(IiSGM),马德里)

专题命中 其他多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09526 2025-10-13 cs.RO 50%

Dynamic Quadrupedal Legged and Aerial Locomotion via Structure Repurposing

Chenghao Wang, Kaushik Venkatesh Krishnamurthy, Shreyansh Pitroda, Adarsh Salagame, Ioannis Mandralis, Eric Sihite, Alireza Ramezani, Morteza Gharib

机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学) Department of Aerospace Engineering, California Institute of Technology(航空航天工程系,加州理工学院)

专题命中 其他多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08636 2025-10-13 astro-ph.IM astro-ph.EP 50%

Foundation Models for Astrobiology: Paper I -- Workshop and Overview

Ryan Felton, Caleb Scharf, Stuart Bartlett, Nathalie A. Cabrol, Victoria Da Poian, Diana Gentry, Jian Gong, Adrienne Hoarfrost, Manil Maskey, Floyd Nichols, Conor A. Nixon, Tejas Panambur, Joseph Pasterski, Anton S. Petrov, Anirudh Prabhu, Brenda Thomson, Hamed Valizadegan, Kimberley Warren-Rhodes, David Wettergreen, Michael L. Wong, Anastasia Yanchilina

专题命中 其他多模态 :multimodal(abstract)

Comments 39 pages, 6 figures, 2 tables, 1 glossary, 4 supplemental pages

详情

展开后加载摘要…

URL PDF HTML 收藏