arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-02 至 2025-10-02 共收录 45 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2510.00855 2025-10-02 cs.CV cs.AI cs.CL cs.LG 67%

Can World Models Benefit VLMs for World Dynamics?

Kevin Zhang, Kuangzhi Ge, Xiaowei Chi, Renrui Zhang, Shaojun Shi, Zhen Dong, Sirui Han, Shanghang Zhang

机构 * Peking University(北京大学) Hong Kong University of Science and Technology(香港科学与技术大学) Chinese University of Hong Kong(香港中文大学) University of California, Santa Barbara(加州大学圣芭芭拉分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project page: https://dyva-worldlm.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00690 2025-10-02 cs.AI 57%

ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning

Yunhao Wang, Ziting Li, Shuai Chen, Tao Liu, Chao Song, Junjie Jiang, Jian Zhu, Peng Gao, Bin Qin

机构 * Xiaomi Inc.(小米公司)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17349 2025-10-02 cs.CV 57%

Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models

Jianing Qi, Jiawei Liu, Hao Tang, Zhigang Zhu

机构 * CUNY Graduate Center(纽约大学研究生中心) Borough of Manhattan Community College(曼哈顿社区学院) The City College of New York(纽约城市学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08561 2025-10-02 cs.CV 57%

ComicsPAP: understanding comic strips by picking the correct panel

Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui, Marco Bertini, Ernest Valveny Llobet, Dimosthenis Karatzas

机构 * CVC, Autonomous University of Barcelona, Spain(CVC,巴塞罗那自治大学,西班牙) MICC, University of Florence, Italy(MICC,佛罗伦萨大学,意大利)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Journal ref Document Analysis and Recognition - ICDAR 2025. ICDAR 2025. Lecture Notes in Computer Science, vol 16023. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2510.00050 2025-10-02 cs.MM cs.AI cs.CV cs.SD eess.AS 83%

Object-AVEdit: An Object-level Audio-Visual Editing Model

Youquan Fu, Ruiyang Si, Hongfa Wang, Dongzhan Zhou, Jiacheng Sun, Ping Luo, Di Hu, Hongyuan Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI)(人工智能研究院) China Telecom(中国电信) Gaoling School of Artificial Intelligence(东城区人工智能学院) Renmin University of China(中国人民大学) Beijing University of Posts and Telecommunications(北京邮电大学) Tencent Data Platform(腾讯数据平台) Tsinghua University(清华大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Huawei Noah’s Ark Lab(华为诺亚实验室) The University of Hong Kong(香港大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00485 2025-10-02 cs.SD cs.AI eess.AS 81%

PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation

Yujia Xiao, Liumeng Xue, Lei He, Xinyi Chen, Aemon Yat Fei Chiu, Wenjie Tian, Shaofei Zhang, Qiuqiang Kong, Xinfa Zhu, Wei Xue, Tan Lee

机构 * The Chinese University of Hong Kong, Hong Kong, China(香港中文大学) The Hong Kong University of Science and Technology, Hong Kong, China(香港科学与技术大学) Microsoft, China(微软公司) South China University of Technology, China(华南理工大学) Northwestern Polytechnical University, China(西北工业大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09746 2025-10-02 cs.SD eess.AS 57%

Deep Learning for Tuberculosis Screening in a High-burden Setting using Cough Analysis and Speech Foundation Models

Ning Ma, Bahman Mirheidari, Guy J. Brown, Nsala Sanjase, Minyoi M. Maimbolwa, Solomon Chifwamba, Seke Muzazu, Monde Muyoyeta, Mary Kagujje

机构 * School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院) Centre for Infectious Disease Research in Zambia(赞比亚传染病疾病研究中心)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments submitted to IEEE Journal of Biomedical and Health Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 6 篇

2506.23972 2025-10-02 cs.CV 83%

Learning Frequency and Memory-Aware Prompts for Multi-Modal Object Tracking

Boyue Xu, Ruichao Hou, Tongwei Ren, Dongming zhou, Gangshan Wu, Jinde Cao

机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) School of Information Science and Engineering(信息科学与工程学院) School of Mathematics(数学学院) Purple Mountain Laboratories(紫金山实验室)

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22646 2025-10-02 cs.CV cs.AI cs.CL 82%

Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs

Xingyu Fu, Siyi Liu, Yinuo Xu, Pan Lu, Guangqiuse Hu, Tianbo Yang, Taran Anantasagar, Christopher Shen, Yikai Mao, Yuanzhe Liu, Keyush Shah, Chung Un Lee, Yejin Choi, James Zou, Dan Roth, Chris Callison-Burch

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://deeptracereward.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25393 2025-10-02 cs.CV cs.AI 81%

Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction

Wendong Yao, Binhua Huang, Soumyabrata Dev

机构 * ADAPT SFI Research Centre, School of Computer Science, University College Dublin(ADAPT SFI研究所以及计算机科学学院,都柏林大学学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments This paper is submitted to IEEE Transactions on Geoscience and Remote Sensing for reviewing

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17481 2025-10-02 eess.SP cs.AI cs.LG 79%

Toward Foundational Model for Sleep Analysis Using a Multimodal Hybrid Self-Supervised Learning Framework

Cheol-Hui Lee, Hakseung Kim, Byung C. Yoon, Dong-Joo Kim

机构 * Department of Brain and Cognitive Engineering, Korea University(脑科学与认知工程系,韩国大学) Interdisciplinary Program in Precision Public Health, Korea University(精准公共卫生跨学科项目,韩国大学) Department of Radiology, Stanford University School of Medicine(放射科,斯坦福大学医学院) VA Palo Alto Health Care System(帕洛阿尔托医疗系统) Department of Neurology, Korea University College of Medicine(神经病学系,韩国大学医学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 18 pages, 5 figures

Journal ref IEEE Transactions on Cybernetics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19778 2025-10-02 cs.AI 79%

Multimodal Large Language Models for Bioimage Analysis

Shanghang Zhang, Gaole Dai, Tiejun Huang, Jianxu Chen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04528 2025-10-02 cs.LG 50%

Federated Dynamic Modeling and Learning for Spatiotemporal Data Forecasting

Thien Pham, Angelo Furno, Faïcel Chamroukhi, Latifa Oukhellou

机构 * COSYS-GRETTIA, Gustave Eiffel University, 77420 France(COSYS-GRETTIA,巴黎-伊夫林大学) ENTPE, University of Lyon(ENTPE,里昂大学) the LICIT-ECO7 University Gustave Eiffel, France(LICIT-ECO7 巴黎-伊夫林大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 1 篇

2505.20152 2025-10-02 cs.CV cs.AI cs.CL 82%

MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models

Kai Sun, Yushi Bai, Zhen Yang, Jiajie Zhang, Ji Qi, Lei Hou, Juanzi Li

机构 * Tsinghua University(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2510.00647 2025-10-02 cs.CL 79%

MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation

Jinlan Fu, Shenzhen Huangfu, Hao Fei, Yichong Huang, Xiaoyu Shen, Xipeng Qiu, See-Kiong Ng

机构 * National University of Singapore(新加坡国立大学) Fudan University(复旦大学) Harbin Institute of Technology(哈尔滨工业大学) Eastern Institute of Technology(东方技术研究所)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CL

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08678 2025-10-02 cs.CV 79%

ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction

Juan Yeo, Soonwoo Cha, Jiwoo Song, Hyunbin Jin, Taesup Kim

专题命中 多模态生成 :any-to-any(title,abstract);分类 cs.CV

Comments Accepted at ICCV25

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05689 2025-10-02 cs.CV 79%

GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving

Zebin Xing, Xingyu Zhang, Yang Hu, Bo Jiang, Tong He, Qian Zhang, Xiaoxiao Long, Wei Yin

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Horizon Robotics Nanjing University(南京大学) Huazhong University of Science & Technology(华中科技大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00046 2025-10-02 cs.CV cs.AI 62%

Reinforcement Learning-Based Prompt Template Stealing for Text-to-Image Models

Xiaotian Zou

机构 * Xiaotian Zou

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00974 2025-10-02 cs.CV 57%

JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation

Siheng Wan, Zhengtao Yao, Zhengdao Li, Junhao Dong, Yanshu Li, Yikai Li, Linshan Li, Haoyan Xu, Yijiang Li, Zhikang Dong, Huacan Wang, Jifeng Shen

机构 * Jiangsu University(江苏大学) University of Southern California(南加州大学) Peking University(北京大学) Nanyang Technological University(南洋理工大学) Brown University(布朗大学) UC San Diego(加州大学圣地亚哥分校) Stony Brook University(石溪大学) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00911 2025-10-02 cs.LG cs.AI 57%

RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training

Tao Ren, Jinyang Jiang, Hui Yang, Wan Tian, Minhao Zou, Guanghao Li, Zishi Zhang, Qinghao Wang, Shentao Qin, Yanjun Zhao, Rui Tao, Hui Shao, Yijie Peng

机构 * Peking University(北京大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00428 2025-10-02 cs.LG cs.AI 57%

Automated Structured Radiology Report Generation with Rich Clinical Context

Seongjae Kang, Dong Bok Lee, Juho Jung, Dongseop Kim, Won Hwa Kim, Sunghoon Joo

机构 * VUNO Inc.(VUNO公司) KAIST(韩国科学技术院) POSTECH(POSTECH大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 34 pages, 30 figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00351 2025-10-02 cs.LG q-bio.BM 50%

Flow Autoencoders are Effective Protein Tokenizers

Rohit Dilip, Evan Zhang, Ayush Varshney, David Van Valen

机构 * California Institute of Technology(加州理工学院) OpenAI(开放人工智能公司) Howard Hughes Medical Institute(霍华德·休斯医学研究院)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 6 篇

2506.05453 2025-10-02 cs.CL cs.AI cs.CV 89%

MLLM-CL: Continual Learning for Multimodal Large Language Models

Hongbo Zhao, Fei Zhu, Haiyang Guo, Meng Wang, Rundong Wang, Gaofeng Meng, Zhaoxiang Zhang

机构 * UCAS(中国科学院自动化研究所) CASIA(中国科学院自动化研究所) HKISI, CAS(中国科学院香港研究所) HKU(香港大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00069 2025-10-02 cs.CV 83%

OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding

Jiancong Xie, Wenjin Wang, Zhuomeng Zhang, Zihan Liu, Qi Liu, Ke Feng, Zixun Sun, Yuedong Yang

机构 * School of Computer Science and Engineering, Sun Yat-sen University, China(中山大学计算机科学与工程学院) Interactive Entertainment Group, Tencent Inc, China(腾讯互动娱乐集团) Key Laboratory of Machine Intelligence and Advanced Computing (MOE), China(机器智能与先进计算重点实验室)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00583 2025-10-02 cs.HC 82%

Rethinking Wine Tasting for Chinese Consumers: A Service Design Approach Enhanced by Multimodal Personalization

Xinyang Shan, Yuanyuan Xu, Tian Xia, Yinshan Lin

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00703 2025-10-02 cs.RO 78%

MultiPhysio-HRC: Multimodal Physiological Signals Dataset for industrial Human-Robot Collaboration

Andrea Bussolan, Stefano Baraldo, Oliver Avram, Pablo Urcola, Luis Montesano, Luca Maria Gambardella, Anna Valente

机构 * ARM-Lab, SUPSI(ARM实验室,SUPSI) Faculty of Informatics, USI(信息学院,USI) Bitbrain Universidad de Zaragoza(萨拉戈萨大学)

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00818 2025-10-02 cs.CV 57%

PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset

Thomas Campagnolo, Ezio Malis, Philippe Martinet, Gaetan Bahl

机构 * Centre Inria d’Universite Cote d’Azur(法国国家信息与自动化技术研究所(Inria))

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted to X-Sense Ego-Exo Sensing for Smart Mobility Workshop at ICCV 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00738 2025-10-02 cs.HC 50%

Datasets for Valence and Arousal Inference: A Survey

Helen Schneider, Svetlana Pavlitska, Helen Gremmelmaier, J. Marius Zöllner

专题命中 多模态评测 :multimodal(abstract)

Comments Accepted for publication at ABAW Workshop at CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 5 篇

2510.00161 2025-10-02 cs.CL 79%

TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Ken Fukuda, Teruko Mitamura

机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所) National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

Comments 21 pages. Code: https://github.com/kimihiroh/tama

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22929 2025-10-02 cs.CL cs.AI cs.CV cs.MA 67%

EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow

Xiaoyu Pan, Yang Bai, Ke Zou, Yang Zhou, Jun Zhou, Huazhu Fu, Yih-Chung Tham, Yong Liu

机构 * Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR)(高性能计算研究所,科技研究局(A*STAR)) Centre for Innovation and Precision Eye Health(创新与精准眼健康中心) Department of Ophthalmology, NUHS Tower Block, Level 7, 1E Kent Ridge Road, Singapore, 119228(眼科部,NUHS塔楼7层,1E Kent Ridge Road,新加坡,119228) Singapore Eye Research Institute, Singapore National Eye Centre, 20 College Road, Singapore, 169856(新加坡眼研究 institute,新加坡国家眼科中心,20 College Road,新加坡,169856)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 9 figures, 5 tables. submit/6621751

详情

展开后加载摘要…

URL PDF HTML 收藏