arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-10 至 2025-09-10 共收录 41 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 10 篇

2505.12363 2025-09-10 cs.CV cs.AI cs.CL cs.LG cs.RO 75%

Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts

Qi Feng

机构 * Kyoto University(京都大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 26 pages, 19 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07588 2025-09-10 cs.CL cs.AI 62%

BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment

Andrey Sakhovskiy, Elena Tutubalina

机构 * AIRI Sber AI ISP RAS Research Center for Trusted AI(俄罗斯科学院信息与系统研究所可信人工智能研究中心)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 1 figure, published in "The 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2025)"

Journal ref Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (2025). Association for Computing Machinery, 1152-1164

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07385 2025-09-10 cs.CV 57%

Parse Graph-Based Visual-Language Interaction for Human Pose Estimation

Shibang Liu, Xuemei Xie, Guangming Shi

机构 * Shibang Liu Xuemei Xie Guangming Shi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00700 2025-09-10 cs.CV 57%

Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision

Raehyuk Jung, Seungjun Yu, Hyunjung Shim

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Link to publicly available codes is added

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16304 2025-09-10 cs.CV 57%

SAMba-UNet: SAM2-Mamba UNet for Cardiac MRI in Medical Robotic Perception

Guohao Huo, Ruiting Dai, Ling Shao, Hao Tang

机构 * University of Electronic Science and Technology of China(电子科技大学) University of Chinese Academy of Sciences(中国科学院大学) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00060 2025-09-10 eess.IV cs.CV 57%

Morphology-based non-rigid registration of coronary computed tomography and intravascular images through virtual catheter path optimization

Karim Kadry, Abhishek Karmakar, Andreas Schuh, Kersten Peterson, Michiel Schaap, David Marlevi, Charles Taylor, Elazer Edelman, Farhad Nezami

机构 * Max L. Olender Andreas Schuh Abhishek Karmakar Kersten Petersen Michiel Schaap David Marlevi Charles Taylor Elazer R. Edelman Farhad R. Nezami

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions in Medical Imaging

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 5 篇

2509.07436 2025-09-10 eess.SP 82%

SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding

Feifan Zhang, Yuyang Du, Yifan Xiang, Xiaoyan Liu, Soung Chang Liew

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07742 2025-09-10 cs.HC cs.AI cs.CV 81%

Enhancing Online Learning by Integrating Biosensors and Multimodal Learning Analytics for Detecting and Predicting Student Behavior: A Review

Alvaro Becerra, Ruth Cobos, Charles Lang

机构 * Department of Computer Science Engineering, Universidad Autónoma de Madrid(马德里自治大学计算机科学工程系) Digital Futures Institute, Teachers College Columbia University(哥伦比亚大学教师学院数字未来研究所)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted for publication in Behaviour & Information Technology (Taylor & Francis). Final published version will be available soon at https://www.tandfonline.com/journals/tbit20

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07422 2025-09-10 eess.SP 78%

Multi-Modal Intelligent Channel Modeling Framework for 6G-Enabled Networked Intelligent Systems

Lu Bai, Zengrui Han, Xuesong Cai, Xiang Cheng

专题命中 其他多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07445 2025-09-10 cs.RO cs.AI 57%

Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions

Harrison Field, Max Yang, Yijiong Lin, Efi Psomopoulou, David Barton, Nathan F. Lepora

机构 * School of Computer Science, University of Bristol, 2 Bristol Robotics Laboratory 3 School of Engineering Mathematics and Technology, University of Bristol(1 计算机科学学院,布里斯托尔大学 2 布里斯托尔机器人实验室 3 工程数学与技术学院,布里斯托尔大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted at CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07863 2025-09-10 cs.HC 50%

NeuroGaze: A Hybrid EEG and Eye-Tracking Brain-Computer Interface for Hands-Free Interaction in Virtual Reality

Kyle Coutray, Wanyea Barbel, Zack Groth, Joseph J LaViola

专题命中 其他多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏