arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-03 至 2025-09-03 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 13 篇

2509.01177 2025-09-03 cs.CV cs.AI cs.HC eess.SP 81%

DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion

Junxiang Liu, Junming Lin, Jiangtong Li, Jie Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00357 2025-09-03 cs.CV cs.AI cs.LG 81%

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Zhen Chen, Xingjian Luo, Kun Yuan, Jinlin Wu, Danny T. M. Chan, Nassir Navab, Hongbin Liu, Zhen Lei, Jiebo Luo

机构 * Hong Kong Institute of Science & Innovation(香港科学与工业创新研究院) CAMP, Technische Universität München(CAMP,慕尼黑技术大学) Department of Surgery, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院外科部)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01338 2025-09-03 cs.AI 79%

Conformal Predictive Monitoring for Multi-Modal Scenarios

Francesca Cairoli, Luca Bortolussi, Jyotirmoy V. Deshmukh, Lars Lindemann, Nicola Paoletti

机构 * University of Trieste, Trieste, Italy(特里埃斯特大学) University of Southern California, Los Angeles, California(南加州大学) King's College London, London, United Kingdom(伦敦国王学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04817 2025-09-03 cs.CV 79%

LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Hongjie Zhang, Lu Dong, Yi Liu, Yifei Huang, Yali Wang, Limin Wang, Yu Qiao

机构 * OpenGVLab, Shanghai AI Laboratory, China(OpenGVLab,上海人工智能实验室) University of Science and Technology of China(中国科学技术大学) Honor Device Co.,Ltd(荣耀设备有限公司) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院) Nanjing University(南京大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15825 2025-09-03 cs.CL q-fin.ST 79%

Enhancing Cryptocurrency Sentiment Analysis with Multimodal Features

Chenghao Liu, Aniket Mahanti, Ranesh Naha, Guanghao Wang, Erwann Sbai

机构 * Department of Computer Science, The University of Auckland(计算机科学系,奥克兰大学) School of Information Systems, Queensland University of Technology(信息系统学院,昆士兰技术大学) Department of Economics, The University of Auckland(经济学系,奥克兰大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01591 2025-09-03 cs.CV 79%

Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval

Adriano Fragomeni, Dima Damen, Michael Wray

机构 * School of Computer Science University of Bristol(计算机科学学院英国布里斯托尔大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05395 2025-09-03 cs.CV cs.IR cs.MM eess.IV 73%

TriPSS: A Tri-Modal Keyframe Extraction Framework Using Perceptual, Structural, and Semantic Representations

Mert Can Cakmak, Nitin Agarwal, Diwash Poudel

机构 * Computer and Information Science, University of Arkansas - Little Rock(计算机与信息科学,亚拉荷马州立大学) ICSI, University of California, Berkeley(ICSI,加州大学伯克利分校) COSMOS Research Center, University of Arkansas - Little Rock(COSMOS研究中心,亚拉荷马州立大学)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13919 2025-09-03 cs.CV cs.AI cs.CL cs.LG cs.RO 67%

Temporal Preference Optimization for Long-Form Video Understanding

Rui Li, Xiaohan Wang, Yuhui Zhang, Orr Zohar, Zeyu Wang, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21496 2025-09-03 cs.CV cs.AI 62%

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

Hao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang, Xuanyu Zheng, Yepeng Tang, Dahua Lin, Lewei Lu

机构 * Sensetime(秒氏科技)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01383 2025-09-03 cs.CV cs.MM 62%

Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning

Long Zhang, Peipei Song, Jianfeng Dong, Kun Li, Xun Yang

机构 * University of Science and Technology of China(中国科学技术大学) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发式智能感知与认知实验室) Zhejiang Gongshang University(浙江工商大学) ReLER, CCAI, Zhejiang University(ReLER,中国计算机学会,浙江大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00210 2025-09-03 cs.CV cs.AI 62%

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment

Jinzhou Tang, Jusheng zhang, Sidi Liu, Waikit Xiu, Qinhan Lv, Xiying Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16834 2025-09-03 cs.LG cs.AI physics.ao-ph 57%

Improving Significant Wave Height Prediction Using Chronos Models

Yilin Zhai, Hongyuan Shi, Chao Zhan, Qing Wang, Zaijin You, Nan Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2403.07815 by other authors

Journal ref Ocean Engineering, Volume 341, Part 2, 1 December 2025, Article 122502

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02353 2025-09-03 cs.RO 50%

SAVOR: Skill Affordance Learning from Visuo-Haptic Perception for Robot-Assisted Bite Acquisition

Zhanxin Wu, Bo Ai, Tom Silver, Tapomayukh Bhattacharjee

机构 * Cornell University(康奈尔大学) UC San Diego(南加州大学)

专题命中 视频多模态 :multi-modal(abstract)

Comments Conference on Robot Learning, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏