arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-21 至 2025-10-21 共收录 94 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 7 篇

2510.16924 2025-10-21 cs.CL 57%

Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?

Zhihui Yang, Yupei Wang, Kaijie Mo, Zhe Zhao, Renfen Hu

机构 * Beijing Normal University(北京师范大学) Tencent AI Lab(腾讯AI实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 (Findings). This version corrects a redundant sentence in the Results section that appeared in the camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16572 2025-10-21 cs.AI cs.MA 57%

Ripple Effect Protocol: Coordinating Agent Populations

Ayush Chopra, Aman Sharma, Feroz Ahmad, Luca Muscariello, Vijoy Pandey, Ramesh Raskar

机构 * Massachusetts Institute of Technology(麻省理工学院) Project Iceberg Cisco(思科)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16457 2025-10-21 cs.CV cs.RO 57%

NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation

Peiran Xu, Xicheng Gong, Yadong MU

机构 * Peking University(北京大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16615 2025-10-21 physics.soc-ph 50%

FlexBSS: A flexible multi-objective framework for bike-sharing station optimization

Jordi Grau-Escolano, David Duran-Rodas, Julian Vicens

专题命中 多模态Agent :multimodal(abstract)

Comments 56 pages, 24 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16524 2025-10-21 cs.RO 50%

Semi-Peaucellier Linkage and Differential Mechanism for Linear Pinching and Self-Adaptive Grasping

Haokai Ding, Zhaohan Chen, Tao Yang, Wenzeng Zhang

专题命中 多模态Agent :multimodal(abstract)

Comments 6 pages, 9 figures, Accepted author manuscript for IEEE CASE 2025

Journal ref 2025 IEEE 21st International Conference on Automation Science and Engineering (CASE), Los Angeles, CA, USA, 2025, pp. 3441-3446

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 18 篇

2510.17205 2025-10-21 cs.CV cs.CL 88%

$\mathcal{V}isi\mathcal{P}runer$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs

Yingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong, Hui Su, Yijie Pan, Wei Zhang, Xiaoyu Shen

机构 * Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT, Ningbo(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT,宁波) Shanghai Jiao Tong University(上海交通大学) Hong Kong Polytechnic University(香港理工大学) Meituan Inc.(美团公司) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15948 2025-10-21 cs.AI cs.CR 85%

VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search

MingSheng Li, Guangze Zhao, Sichen Liu

机构 * Independent Researcher(独立研究者) Harbin Institute of Technology(哈尔滨工业大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16824 2025-10-21 cs.LG q-bio.MN 82%

ProtoMol: Enhancing Molecular Property Prediction via Prototype-Guided Multimodal Learning

Yingxu Wang, Kunyu Zhang, Jiaxin Huang, Nan Yin, Siwei Liu, Eran Segal

机构 * MBZUAI(穆扎芬人工智能研究所) University of Zhengzhou(郑州大学) HKUST(香港科技大学) University of Aberdeen(爱丁堡大学) MBZUAI, Weizmann Institute of Science(穆扎芬人工智能研究所、威斯曼科学研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16350 2025-10-21 cs.LG 82%

MGTS-Net: Exploring Graph-Enhanced Multimodal Fusion for Augmented Time Series Forecasting

Shule Hao, Junpeng Bao, Wenli Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00916 2025-10-21 cs.CV cs.AI 81%

Enhancing Osteoporosis Detection: An Explainable Multi-Modal Learning Framework with Feature Fusion and Variable Clustering

Mehdi Hosseini Chagahi, Saeed Mohammadi Dashtaki, Niloufar Delfan, Nadia Mohammadi, Farshid Rostami Pouria, Behzad Moshiri, Md. Jalil Piran, Oliver Faust

机构 * School of Electrical and Computer Engineering, College of Engineering, University of Tehran(塔里班大学电气与计算机工程学院) Department of Epidemiology, Shiraz University of Medical Science(谢尔兹医学科学大学流行病学系) Department of Computer Science and Engineering, Sejong University(世宗大学计算机科学与工程系) School of Computing and Information Science, Anglia Ruskin University(安格利亚 Ruskin 大学计算与信息科学学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17394 2025-10-21 cs.LG cs.CV 79%

MILES: Modality-Informed Learning Rate Scheduler for Balancing Multimodal Learning

Alejandro Guerra-Manzanares, Farah E. Shamout

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted and presented at the 2025 International Joint Conference on Neural Networks (IJCNN'25). The paper was awarded an honorable mention (best 4 papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17289 2025-10-21 cs.CL 79%

Addressing Antisocial Behavior in Multi-Party Dialogs Through Multimodal Representation Learning

Hajar Bakarou, Mohamed Sinane El Messoussi, Anaïs Ollagnier

机构 * Universit\'e C \ te d'Azur, CNRS, Inria, I3S Sophia Antipolis France Universit\'e C \ te d'Azur, CNRS, Inria, I3S

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17078 2025-10-21 cs.CV 79%

Towards a Generalizable Fusion Architecture for Multimodal Object Detection

Jad Berjawi, Yoann Dupas, Christophe C'erin

机构 * Université Grenoble Alpes(格勒诺布尔大学) Université Sorbonne Paris Nord(巴黎-萨克勒大学) INRIA(法国国家信息与自动化研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 8 figures, accepted at ICCV 2025 MIRA Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08217 2025-10-21 cs.AI 79%

Quantum Federated Learning for Multimodal Data: A Modality-Agnostic Approach

Atit Pokharel, Ratun Rahman, Thomas Morris, Dinh C. Nguyen

机构 * Department of Electrical and Computer Engineering, The University of Alabama in Huntsville(电气与计算机工程系,阿拉巴马大学亨茨维尔分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments This paper was presented at BEAM with CVPR 2025

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 545-554. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15944 2025-10-21 cs.LG cs.AI 79%

Lyapunov-Stable Adaptive Control for Multimodal Concept Drift

Tianyu Bell Pan, Mengdi Zhu, Alexa Jordyn Cole, Ronald Wilson, Damon L. Woodard

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Florida Institute of National Security(佛罗里达国家安全研究所) Applied Artificial Intelligence Group(应用人工智能组) University of Florida(佛罗里达大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21514 2025-10-21 cs.CV 79%

G$^{2}$D: Boosting Multimodal Learning with Gradient-Guided Distillation

Mohammed Rakib, Arunkumar Bagavathi

机构 * Oklahoma State University(俄克拉荷马州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16086 2025-10-21 cs.LG stat.AP 78%

FSRF: Factorization-guided Semantic Recovery for Incomplete Multimodal Sentiment Analysis

Ziyang Liu, Pengjunfei Chu, Shuming Dong, Chen Zhang, Mingcheng Li, Jin Wang

机构 * School of Future Science and Engineering, Soochow University(苏州大学未来科学与工程学院) School of Advanced Manufacturing Engineering, Hefei University(合肥大学先进制造工程学院) College of Global Talents, BITZH, Beijing Institute of Technology(北京理工大学珠海学院全球人才学院) Academy for Engineering and Technology, Fudan University(复旦大学工程与技术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 6 pages,3 figures

Journal ref In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15953 2025-10-21 cs.CR 78%

Hierarchical Multi-Modal Threat Intelligence Fusion Without Aligned Data: A Practical Framework for Real-World Security Operations

Sisir Doppalapudi

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01719 2025-10-21 cs.LG 78%

Robust Anomaly Detection through Multi-Modal Autoencoder Fusion for Small Vehicle Damage Detection

Sara Khan, Mehmed Yüksel, Frank Kirchner

机构 * Faculty of Mathematics and Computer Science, University of Bremen(数学与计算机科学学院,不莱梅大学) Robotics Innovation Center, Deutsches Forschungszentrum für Künstliche Intelligenz(机器人创新中心,德国人工智能研究中心) Engineering Software Communication, Robert Bosch GmbH(工程软件通信,罗伯特·博世有限公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments 17 pages, 12 figures, submitted to Elsevier MLWA

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04863 2025-10-21 cs.LG q-bio.BM 78%

OneProt: Towards Multi-Modal Protein Foundation Models

Klemens Flöge, Srisruthi Udayakumar, Johanna Sommer, Marie Piraud, Stefan Kesselheim, Vincent Fortuin, Stephan Günneman, Karel J van der Weg, Holger Gohlke, Erinc Merdivan, Alina Bazarova

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments 34 pages, 7 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17034 2025-10-21 cs.CV 57%

Where, Not What: Compelling Video LLMs to Learn Geometric Causality for 3D-Grounding

Yutong Zhong

机构 * New York University(纽约大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10351 2025-10-21 cs.LG cs.AI 57%

PhysioWave: A Multi-Scale Wavelet-Transformer for Physiological Signal Representation

Yanlong Chen, Mattia Orlandi, Pierangelo Maria Rapa, Simone Benatti, Luca Benini, Yawei Li

机构 * IIS, ETH Zurich(苏黎世联邦理工学院智能信息系统实验室) DEI, University of Bologna(博洛尼亚大学电子工程学院) DIEF, University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学经济工程学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments 43 pages, 17 figures, 17 tables. Accepted by NeurIPS 2025. Code and data are available at: github.com/ForeverBlue816/PhysioWave

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16555 2025-10-21 cs.AI cs.LG 57%

Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence

Qiongyan Wang, Xingchen Zou, Yutian Jiang, Haomin Wen, Jiaheng Wei, Qingsong Wen, Yuxuan Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Carnegie Mellon University(卡内基梅隆大学) Squirrel Ai Learning

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 11 篇

2510.16785 2025-10-21 cs.CV 83%

Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs

Jiazhen Liu, Long Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12868 2025-10-21 cs.CV 79%

Computer-Aided Design of Personalized Occlusal Positioning Splints Using Multimodal 3D Data

Agnieszka Anna Tomaka, Leszek Luchowski, Michał Tarnawski, Dariusz Pojda

机构 * Institute of Theoretical and Applied Informatics, Polish Academy of Sciences(理论与应用信息学研究所,波兰科学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14022 2025-10-21 eess.IV cs.CV 79%

I2I-Mamba: Multi-modal medical image synthesis via selective state space modeling

Omer F. Atli, Bilal Kabas, Fuat Arslan, Arda C. Demirtas, Mahmut Yurt, Onat Dalmaz, Tolga Çukur

机构 * Department of Electrical and Electronics Engineering, and National Magnetic Resonance Research Center, Bilkent University(电子工程系和国家磁共振研究中心,比尔肯特大学) Department of Electrical Engineering, Stanford University(电气工程系,斯坦福大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17686 2025-10-21 cs.CV 70%

Towards 3D Objectness Learning in an Open World

Taichi Liu, Zhenyu Wang, Ruofeng Liu, Guang Wang, Desheng Zhang

机构 * Rutgers University(罗格斯大学) Tsinghua University(清华大学) Michigan State University(密歇根州立大学) Florida State University(佛罗里达州立大学)

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07878 2025-10-21 cs.LG cs.AI eess.SP q-bio.NC 61%

Comparative Analysis of Deep Learning Approaches for Harmful Brain Activity Detection Using EEG

Shivraj Singh Bhatti, Aryan Yadav, Mitali Monga, Neeraj Kumar

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Thapar Institute of Engineering and Technology(泰帕尔工程与技术学院)

专题命中 其他多模态 :multimodal(abstract,comments);分类 cs.AI

Comments 6 pages, 5 figures. Presented at IEEE CICT 2024. The paper discusses the application of multimodal data and training strategies in EEG-based brain activity classification

Journal ref 2024 IEEE 8th Int. Conf. on Info. and Comm. Tech. (CICT), pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16034 2025-10-21 cs.CV 57%

VisualLens: Personalization through Task-Agnostic Visual History

Wang Bill Zhu, Deqing Fu, Kai Sun, Yi Lu, Zhaojiang Lin, Seungwhan Moon, Kanika Narang, Mustafa Canim, Yue Liu, Anuj Kumar, Xin Luna Dong

机构 * Meta University of Southern California(南加州大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17393 2025-10-21 q-fin.PM 50%

3S-Trader: A Multi-LLM Framework for Adaptive Stock Scoring, Strategy, and Selection in Portfolio Optimization

Kefan Chen, Hussain Ahmad, Diksha Goel, Claudia Szabo

专题命中 其他多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏