arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-13 至 2025-11-13 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 5 篇

2505.23990 2025-11-13 cs.AI 79%

Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding

Mingyang Mao, Mariela M. Perez-Cabarcas, Utteja Kallakuri, Nicholas R. Waytowich, Xiaomin Lin, Tinoosh Mohsenin

机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09342 2025-11-13 eess.SP 78%

A cross-modal pre-training framework with video data for improving performance and generalization of distributed acoustic sensing

Junyi Duan, Jiageng Chen, Zuyuan He

专题命中 视频多模态 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09137 2025-11-13 eess.SP 78%

xHAP: Cross-Modal Attention for Haptic Feedback Estimation in the Tactile Internet

Georgios Kokkinis, Alexandros Iosifidis, Qi Zhang

专题命中 视频多模态 :cross-modal(title,abstract)

Comments 12 pages, 13 figures, 3 tables, 2 algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03184 2025-11-13 eess.IV cs.CV 57%

EvRWKV: A Continuous Interactive RWKV Framework for Effective Event-Guided Low-Light Image Enhancement

Wenjie Cai, Qingguo Meng, Zhenyu Wang, Xingbo Dong, Zhe Jin

机构 * Anhui Provincial International Joint Research center for Advanced technology in Medical imaging(安徽省国际联合先进医学影像技术研究中心) School of Artificial Intelligence(人工智能学院) Anhui University(安徽大学) State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology(光电信息采集与防护技术国家重点实验室) Anhui Provincial Key Laboratory of Secure Artificial Intelligence(安徽省安全人工智能重点实验室) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(粤港澳大湾区人工智能与数字经济实验室(深圳))

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08892 2025-11-13 cs.HC cs.RO 50%

Help or Hindrance: Understanding the Impact of Robot Communication in Action Teams

Tauhid Tanjim, Jonathan St. George, Kevin Ching, Angelique Taylor

机构 * Department of Information Science at Cornell University(康奈尔大学信息科学系) Weill Cornell Medicine, Cornell University(韦尔医学院,康奈尔大学)

专题命中 视频多模态 :multimodal(abstract)

Comments This is the author's original submitted version of the paper accepted to the 2025 IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). \c{opyright} 2025 IEEE. Personal use of this material is permitted. For any other use, please contact IEEE

Journal ref 2025 34th IEEE International Conference on Robot and Human Interactive Communication RO-MAN pp. 1460-1465

详情

展开后加载摘要…

URL PDF HTML 收藏