arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-17 至 2025-09-17 共收录 53 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 12 篇

2509.12540 2025-09-17 cs.LG 78%

Cross-Modal Deep Metric Learning for Time Series Anomaly Detection

Wei Li, Zheze Yang

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 多模态评测 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12243 2025-09-17 math.NA cs.NA 78%

Bi-fidelity Interpolative Decomposition for Multimodal Data

Murray Cutforth, Tiffany Fan, Tony Zahtila, Alireza Doostan, Eric Darve

专题命中 多模态评测 :multimodal(title);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05066 2025-09-17 cs.CL cs.AI 62%

ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions

Matteo Bortoletto, Constantin Ruhdorfer, Andreas Bulling

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07825 2025-09-17 cs.CV 57%

3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark

Wufei Ma, Haoyu Chen, Guofeng Zhang, Yu-Cheng Chou, Jieneng Chen, Celso M de Melo, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) Carnegie Mellon University(卡内基梅隆大学) DEVCOM Army Research Laboratory(DEVCOM陆军研究实验室)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments ICCV 2025. Project page: https://3dsrbench.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12816 2025-09-17 cs.HC cs.AI cs.CV cs.LG 54%

Gesture Evaluation in Virtual Reality

Axel Wiebe Werner, Jonas Beskow, Anna Deichler

机构 * Division of Speech, Music and Hearing, KTH Royal Institute of Technology(语音、音乐与听觉系,皇家理工学院)

专题命中 多模态评测 :multimodal(comments,journal_ref);分类 cs.CV、cs.AI

Comments Published in Proceedings of the 26th International Conference on Multimodal Interaction (ICMI '24), ACM. Copyright 2024 ACM. Licensed under CC BY

Journal ref Proceedings of the 26th International Conference on Multimodal Interaction (ICMI '24), ACM, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 3 篇

2509.12279 2025-09-17 cs.CV cs.AI 62%

Domain Adaptive SAR Wake Detection: Leveraging Similarity Filtering and Memory Guidance

He Gao, Baoxiang Huang, Milena Radenkovic, Borui Li, Ge Chen

机构 * College of Computer Science and Technology, Qingdao University(青岛大学计算机科学与技术学院) Laboratory for Regional Oceanography and Numerical Modeling, Qingdao Marine Science and Technology Center(青岛海洋科学与技术中心区域海洋学与数值模拟实验室) School of Computer Science and Information Technology, The University of Nottingham(诺丁汉大学计算机科学与信息学院) State Key Laboratory of Physical Oceanography, Department of Marine Technology, Ocean University of China(中国海洋大学物理海洋学国家重点实验室,海洋技术系)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12618 2025-09-17 cs.RO cs.AI cs.CV 62%

ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation

Zekai Zhang, Weiye Zhu, Hewei Pan, Xiangchen Wang, Rongtao Xu, Xing Sun, Feng Zheng

机构 * Southern University of Science and Technology, China(南方科技大学,中国) Spatialtemporal AI, China(时空AI,中国) MBZUAI, UAE(MBZUAI,阿联酋) Tencent Youtu Lab(腾讯优图实验室)

专题命中 多模态Agent :MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12423 2025-09-17 cs.AI cs.CL 62%

Small Models, Big Results: Achieving Superior Intent Extraction through Decomposition

Danielle Cohen, Yoni Halpern, Noam Kahlon, Joel Oren, Omri Berkovitch, Sapir Caduri, Ido Dagan, Anatoly Efros

机构 * Google(谷歌) Bar-Ilan University(巴伊兰大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 10 篇

2509.13070 2025-09-17 cs.CV cs.AI 86%

TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation

Qianqi Lu, Yuxiang Xie, Jing Zhang, Shiwei Zou, Yan Chen, Xidao Luan

专题命中 多模态训练与对齐 :image-text(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12234 2025-09-17 cs.LG cs.AI cs.CV eess.IV 81%

Flexible Multimodal Neuroimaging Fusion for Alzheimer's Disease Progression Prediction

Benjamin Burns, Yuan Xue, Douglas W. Scharre, Xia Ning

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Department of Biomedical Informatics(生物医学信息学系) Department of Neurology(神经病学系) Translational Data Analytics Institute(转化数据分析研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at Applications of Medical AI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12901 2025-09-17 cs.CV 79%

MSGFusion: Multimodal Scene Graph-Guided Infrared and Visible Image Fusion

Guihui Li, Bowei Dong, Kaizhi Dong, Jiayi Li, Haiyong Zheng

机构 * College of Computer Science and Technology, Ocean University of China(中国海洋大学计算机科学与技术学院) College of Electronic Engineering, Ocean University of China(中国海洋大学电子工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05300 2025-09-17 eess.SP 78%

Harnessing Multimodal Sensing for Multi-user Beamforming in mmWave Systems

Kartik Patel, Robert W. Heath

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Journal ref IEEE Trans.Wireless Commun. 23 (2024) 18725-18739

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12930 2025-09-17 cs.DC 78%

Analysis and Optimization of Wireless Multimodal Federated Learning on Modal Heterogeneity

Xuefeng Han, Wen Chen, Jun Li, Ming Ding, Qingqing Wu, Kang Wei, Xiumei Deng, Yumeng Shao, Qiong Wu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10720 2025-09-17 cs.CY 78%

Adapting Public Personas: A Multimodal Study of U.S. Legislators' Cross-Platform Social Media Strategies

Weihong Qi, Anushka Dave, Chen Ling

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19439 2025-09-17 cs.CV 70%

AMF-MedIT: An Efficient Align-Modulation-Fusion Framework for Medical Image-Tabular Data

Congjing Yu, Jing Ye, Yang Liu, Xiaodong Zhang, Zhiyong Zhang

机构 * School of Electronics and Communication Engineering(电子与通信工程学院) Guangdong Provincial Key Laboratory of Advanced IntelliSense Technology(广东省先进智能感知技术重点实验室) Department of Radiology(放射科)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12886 2025-09-17 cs.CL cs.AI 62%

The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations

Yubo Zhu, Dongrui Liu, Zecheng Lin, Wei Tong, Sheng Zhong, Jing Shao

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Xidian University(西安电子科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21886 2025-09-17 cs.AI cs.LG eess.SP 57%

Efficient Pain Recognition via Respiration Signals: A Single Cross-Attention Transformer Multi-Window Fusion Pipeline

Stefanos Gkikas, Ioannis Kyprakis, Manolis Tsiknakis

机构 * Foundation for Research \& Technology-Hellas Heraklion Greece Foundation for Research \& Technology-Hellas Hellenic Mediterranean University Heraklion Greece Hellenic Mediterranean University

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2507.21881, arXiv:2507.21875

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21881 2025-09-17 cs.AI 57%

Multi-Representation Diagrams for Pain Recognition: Integrating Various Electrodermal Activity Signals into a Single Image

Stefanos Gkikas, Ioannis Kyprakis, Manolis Tsiknakis

机构 * Foundation for Research \& Technology-Hellas Heraklion Greece Foundation for Research \& Technology-Hellas Hellenic Mediterranean University Heraklion Greece Hellenic Mediterranean University

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2507.21875

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 5 篇

2509.12521 2025-09-17 cs.LG 82%

Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time

Yifan Lan, Yuanpu Cao, Weitong Zhang, Lu Lin, Jinghui Chen

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12620 2025-09-17 cs.RO 78%

PerchMobi^3: A Multi-Modal Robot with Power-Reuse Quad-Fan Mechanism for Air-Ground-Wall Locomotion

Yikai Chen, Zhi Zheng, Jin Wang, Bingye He, Xiangyu Xu, Jialu Zhang, Huan Yu, Guodong Lu

机构 * Zhejiang University(浙江大学) Robotics Research Center of Yuyao City(余姚市机器人研究所) State Key Laboratory of Fluid Power and Mechatronic Systems(流体动力与机电系统国家重点实验室) Zhejiang Key Laboratory of Industrial Big Data and Robot Intelligent Systems(浙江省工业大数据与机器人智能系统重点实验室) College of Control Science and Engineering(控制科学与工程学院)

专题命中 其他多模态 :multi-modal(title,abstract)

Comments 7 pages, 8 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13064 2025-09-17 cs.HC 50%

Patient Perspectives on Telemonitoring during Colorectal Cancer Surgery Prehabilitation

Irina Bianca Serban, Dimitra Dritsa, David ten Cate, Loes Janssen, Margot Heijmans, Sara Colombo, Aarnout Brombacher, Steven Houben

专题命中 其他多模态 :multimodal(abstract)

Comments 20 pages, 3 figures, presented at the 19th EAI International Conference on Pervasive Computing Technologies for Healthcare, to be published in the Springer - LNICST series

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12522 2025-09-17 eess.SY cs.SY 50%

Hybrid State Estimation of Uncertain Nonlinear Dynamics Using Neural Processes

Devin Hunter, Chinwendu Enyioha

专题命中 其他多模态 :multimodal(abstract)

Comments 32 pages (single column) - 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01957 2025-09-17 cs.NI 50%

Federated Foundation Models in Harsh Wireless Environments: Prospects, Challenges, and Future Directions

Evan Chen, Seyyedali Hosseinalipour, Christopher G. Brinton, David J. Love

专题命中 其他多模态 :multimodal(abstract)

Comments This paper is under review in IEEE Network Magazine Special Issue on Large AI Models for the Internet of Everything

详情

展开后加载摘要…

URL PDF HTML 收藏