arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-05 至 2025-08-05 共收录 106 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 18 篇

2406.09016 2025-08-05 cs.CV 79%

Cross-Modal Learning for Anomaly Detection in Complex Industrial Process: Methodology and Benchmark

Gaochang Wu, Yapeng Zhang, Lan Deng, Jingxin Zhang, Tianyou Chai

机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang 110819, P. R. China(合成过程工业系统自动化国家重点实验室,东北大学,沈阳110819,中华人民共和国) School of Software and Electrical Engineering, Swinburne University of Technology, Melbourne, VIC 3122, Australia(软件与电子工程学院,斯威本科技大学,墨尔本,VIC 3122,澳大利亚)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments 14 pages, 6 figures, 5 tables. IEEE TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12937 2025-08-05 cs.AI cs.CL cs.CV cs.LG 78%

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Jingyi Zhang, Jiaxing Huang, Huanjin Yao, Shunyu Liu, Xikun Zhang, Shijian Lu, Dacheng Tao

机构 * Nanyang Technological University(南洋理工大学)

专题命中 多模态评测 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments ICCV 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00974 2025-08-05 cs.CV cs.AI cs.LG 62%

ThermoCycleNet: Stereo-based Thermogram Labeling for Model Transition to Cycling

Daniel Andrés López, Vincent Weber, Severin Zentgraf, Barlo Hillen, Perikles Simon, Elmar Schömer

机构 * Institute of Computer Science, Research Group Computational Geometry, Johannes Gutenberg University, Mainz, Germany(计算机科学研究所,计算几何研究组,吉森大学,马因茨,德国) Institute of Sports Science, Department of Sports Medicine, Disease Prevention and Rehabilitation, Johannes Gutenberg University, Mainz, Germany(体育科学研究所,运动医学部门,疾病预防与康复,吉森大学,马因茨,德国) Institute of Occupational, Social and Environmental Medicine, University Medical Center, Mainz, Germany(职业、社会和环境医学研究所,大学医学中心,马因茨,德国)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Presented at IWANN 2025 18th International Work-Conference on Artificial Neural Networks, A Coruña, Spain, 16-18 June, 2025. Book of abstracts: ISBN: 979-13-8752213-1. Funding: Johannes Gutenberg University "Stufe I'': "Start ThermoCycleNet''. Partial funding: Carl-Zeiss-Stiftung: "Multi-dimensionAI'' (CZS-Project number: P2022-08-010)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00893 2025-08-05 cs.CL cs.AI 62%

Affordance Benchmark for MLLMs

Junying Wang, Wenzhe Li, Yalun Wu, Yingji Liang, Yijin Guo, Chunyi Li, Haodong Duan, Zicheng Zhang, Guangtao Zhai

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14240 2025-08-05 cs.CV cs.AI 62%

CityNav: A Large-Scale Dataset for Real-World Aerial Navigation

Jungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Yutaka Matsuo, Nakamasa Inoue

机构 * Institute of Science Tokyo(东京科学研究所) The University of Tokyo(东京大学) ATR(ATR公司) National Institute of Informatics(信息研究院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments ICCV2025. The first two authors are equally contributed. Project page: https://water-cookie.github.io/city-nav-proj/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02645 2025-08-05 cs.CV 61%

Evaluating Variance in Visual Question Answering Benchmarks

Nikitha SR

机构 * Media and Data Science Research Lab, Adobe(媒体与数据科学研究实验室,Adobe)

专题命中 多模态评测 :multimodal(abstract,comments);分类 cs.CV

Comments Accepted in ICCV 2025 Workshop on What's Next in Multimodal Foundational Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02307 2025-08-05 cs.CV cs.LG 57%

Whole-body Representation Learning For Competing Preclinical Disease Risk Assessment

Dmitrii Seletkov, Sophie Starck, Ayhan Can Erdur, Yundi Zhang, Daniel Rueckert, Rickmer Braren

机构 * Institute of Diagnostic Interventional Radiology, Technical University of Munich, School of Medicine, Munich, Germany Chair for AI in Healthcare Medicine, Technical University of Munich (TUM) TUM University Hospital, Munich, Germany Department of Computing, Imperial College London, London, UK German Cancer Consortium (DKTK), Munich partner site, Heidelberg, Germany

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01704 2025-08-05 cs.CV 57%

LT-Gaussian: Long-Term Map Update Using 3D Gaussian Splatting for Autonomous Driving

Luqi Cheng, Zhangshuo Qi, Zijie Zhou, Chao Lu, Guangming Xiong

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted by IV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01016 2025-08-05 eess.IV cs.CV 57%

Diagnostic Accuracy of Open-Source Vision-Language Models on Diverse Medical Imaging Tasks

Gustav Müller-Franzes, Debora Jutz, Jakob Nikolas Kather, Christiane Kuhl, Sven Nebelung, Daniel Truhn

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18277 2025-08-05 cs.CV cs.LG 57%

Towards Modality Generalization: A Benchmark and Prospective Analysis

Xiaohao Liu, Xiaobo Xia, Zhuo Huang, See-Kiong Ng, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) The University of Sydney(悉尼大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025 (CR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02342 2025-08-05 cs.IR 50%

Agentic Personalized Fashion Recommendation in the Age of Generative AI: Challenges, Opportunities, and Evaluation

Yashar Deldjoo, Nima Rafiee, Mahdyar Ravanbakhsh

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11884 2025-08-05 cs.LG 50%

Out-of-Distribution Detection: A Task-Oriented Survey of Recent Advances

Shuo Lu, Yingsheng Wang, Lijun Sheng, Lingxiao He, Aihua Zheng, Jian Liang

机构 * NLPR \& MAIS, Institute of Automation, Chinese Academy of Sciences Anhui University Hefei China University of Science Anhui University

专题命中 多模态评测 :multi-modal(abstract)

Comments Accepted to ACM Computing Surveys (CSUR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 4 篇

2508.01917 2025-08-05 cs.RO cs.AI 57%

L3M+P: Lifelong Planning with Large Language Models

Krish Agarwal, Yuqian Jiang, Jiaheng Hu, Bo Liu, Peter Stone

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01186 2025-08-05 cs.AI cs.HC 57%

A Survey on Agent Workflow -- Status and Future

Chaojia Yu, Zihan Cheng, Hanwen Cui, Yishuo Gao, Zexu Luo, Yijin Wang, Hangbin Zheng, Yong Zhao

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 12 pages, 3 figures, accepted to IEEE Conference, ICAIBD(International Conference of Artificial Intelligence and Big Data) 2025. This is the author's version, not the publisher's. See https://ieeexplore.ieee.org/document/11082076

Journal ref IEEE ICAIBD 2025, pp. 770-781

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12974 2025-08-05 cs.CV cs.RO 57%

Exploring 3D Reasoning-Driven Planning: From Implicit Human Intentions to Route-Aware Activity Planning

Xueying Jiang, Wenhao Li, Xiaoqin Zhang, Ling Shao, Shijian Lu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17540 2025-08-05 cs.RO cs.LG 50%

EqDrive: Efficient Equivariant Motion Forecasting with Multi-Modality for Autonomous Driving

Yuping Wang, Jier Chen

机构 * University of Michigan(密歇根大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态Agent :multi-modal(abstract)

Comments 6 pages, 7 figures, Accepted 2024 International Conference on Robotics and Automation

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 21 篇

2508.02479 2025-08-05 cs.CV 85%

Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding

Xinquan Yu, Wei Lu, Xiangyang Luo

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02525 2025-08-05 cs.AI 83%

Accurate and Interpretable Postmenstrual Age Prediction via Multimodal Large Language Model

Qifan Chen, Jin Cui, Cindy Duan, Yushuo Han, Yifei Shi

机构 * King’s College London(伦敦国王学院) Imperial College London(帝国理工学院) Columbia University(哥伦比亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Submitted to the NeurIPS 2025 Workshop GenAI4Health. Conference website: https://aihealth.ischool.utexas.edu/GenAI4HealthNeurips2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01644 2025-08-05 cs.MM cs.AI cs.CV cs.SD eess.AS 83%

DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition

Peiyuan Jiang, Yao Liu, Qiao Liu, Zongshun Zhang, Jiaye Yang, Lu Liu, Daibing Yao

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Published in ACM Multimedia 2025. 10 pages, 4 figures

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01805 2025-08-05 cs.NI 82%

M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks

Yongjie Zeng, Hongyang Du

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00926 2025-08-05 cs.LG 82%

Hybrid Hypergraph Networks for Multimodal Sequence Data Classification

Feng Xu, Hui Wang, Yuting Huang, Danwei Zhang, Zizhu Fan

机构 * Feng Xu 1,2(作者1单位) Hui Wang 1(作者1单位) Yuting Huang 3(作者3单位) Danwei Zhang 4(作者4单位) Zizhu Fan 5(作者5单位)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01525 2025-08-05 cs.CV cs.AI 81%

MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection

Kuo Shi, Jie Lu, Shanshan Ye, Guangquan Zhang, Zhen Fang

机构 * University of Technology Sydney(悉尼技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01316 2025-08-05 cs.CV cs.HC 79%

Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust

Mohsen Abbaspour Onari, Lucie Charlotte Magister, Yaoxin Wu, Amalia Lupi, Dario Creazzo, Mattia Tordin, Luigi Di Donatantonio, Emilio Quaia, Chao Zhang, Isel Grau, Marco S. Nobile, Yingqian Zhang, Pietro Liò

机构 * Information Systems Group, Eindhoven University of Technology, The Netherlands(埃因霍温技术大学信息系统组) Eindhoven Artificial Intelligence Systems Institute, The Netherlands(埃因霍温人工智能系统研究所) Department of Computer Science and Technology, University of Cambridge, United Kingdom(剑桥大学计算机科学与技术系) Department of Medicine - DIMED, Padua University Hospital, Italy(帕多瓦大学医院医学部-DIMED) Department of Environmental Sciences, Informatics, and Statistics, Ca’ Foscari University of Venice, Italy(威尼斯卡弗里大学环境科学、信息学与统计学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00963 2025-08-05 cs.LG cs.AI 79%

Rethinking Multimodality: Optimizing Multimodal Deep Learning for Biomedical Signal Classification

Timothy Oladunni, Alex Wong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00830 2025-08-05 cs.CV 79%

Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Yang Cao, Yihan Zeng, Hang Xu, Dan Xu

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科技大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Code Page: NeurIPS2023" target="_blank" rel="noopener">https://github.com/yangcaoai/CoDA_NeurIPS2023 This paper is accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.07348 2025-08-05 cs.CV 79%

Juggling With Representations: On the Information Transfer Between Imagery, Point Clouds, and Meshes for Multi-Modal Semantics

Dominik Laupheimer, Norbert Haala

机构 * Institute for Photogrammetry, University of Stuttgart, Germany(摄影测量研究所,斯图加特大学,德国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02409 2025-08-05 cs.CV cs.AI 76%

Hydra: Accurate Multi-Modal Leaf Wetness Sensing with mm-Wave and Camera Fusion

Yimeng Liu, Maolin Gan, Huaili Zeng, Li Liu, Younsuk Dong, Zhichao Cao

机构 * Michigan State University(密歇根州立大学) Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments In Proceedings of ACM MobiCom (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01150 2025-08-05 cs.CV 70%

OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding

Dianyi Yang, Xihan Wang, Yu Gao, Shiyang Liu, Bohan Ren, Yufeng Yue, Yi Yang

机构 * School of Automation, Beijing Institute of Technology, Beijing, China(自动化学院,北京理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments IROS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02331 2025-08-05 cs.CV cs.SD 70%

VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection

Hao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu, Jia Li, Meng Wang, Richang Hong

机构 * Hefei University of Technology(合肥工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Source code and pre-trained models will be available at https://github.com/MSA-LMC/VAEmo

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19903 2025-08-05 cs.CV 70%

Scaling Vision Pre-Training to 4K Resolution

Baifeng Shi, Boyi Li, Han Cai, Yao Lu, Sifei Liu, Marco Pavone, Jan Kautz, Song Han, Trevor Darrell, Pavlo Molchanov, Hongxu Yin

机构 * UC Berkeley(加州大学伯克利分校) NVIDIA(英伟达)

专题命中 多模态训练与对齐 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2025. Project Page: https://nvlabs.github.io/PS3

详情

展开后加载摘要…

URL PDF HTML 收藏