arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-29 至 2025-07-29 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 16 篇

2507.20451 2025-07-29 cs.AI 79%

STARN-GAT: A Multi-Modal Spatio-Temporal Graph Attention Network for Accident Severity Prediction

Pritom Ray Nobin, Imran Ahammad Rifat

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20163 2025-07-29 cs.CV 79%

Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning

Zeyu Xi, Haoying Sun, Yaofei Wu, Junchi Yan, Haoran Zhang, Lifang Wu, Liang Wang, Changwen Chen

机构 * Beijing University of Technology(北京理工大学) Shanghai Jiao Tong University(上海交通大学) Chinese Academy of Sciences(中国科学院) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025 (Poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10503 2025-07-29 cs.CV cs.CL cs.LG 79%

Everything is a Video: Unifying Modalities through Next-Frame Prediction

G. Thomas Hudson, Dean Slack, Thomas Winterbottom, Jamie Sterling, Chenghao Xiao, Junjie Shentu, Noura Al Moubayed

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13092 2025-07-29 cs.CV 70%

EventVAD: Training-Free Event-Aware Video Anomaly Detection

Yihua Shao, Haojin He, Sijie Li, Siyu Chen, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma, Xiaochen Wang, Hao Tang, Yan Wang, Shuyan Li

机构 * Peking University(北京大学) Guangdong University of Technology(广东工业大学) The University of Sheffield(谢菲尔德大学) University of Science and Technology Beijing(北京科技大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Nanjing University(南京大学) University of Trento(特伦特大学) Queen's University Belfast(贝尔法斯特女王大学)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Paper was accepted by ACM MM 2025; Code: https://github.com/YihuaJerry/EventVAD

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17402 2025-07-29 cs.CV cs.IR cs.MM 62%

HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

Jun Li, Jinpeng Wang, Chaolei Tan, Niu Lian, Long Chen, Yaowei Wang, Min Zhang, Shu-Tao Xia, Bin Chen

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Research Center of Artificial Intelligence, Peng Cheng Laboratory(鹏城实验室人工智能研究中心) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by ICCV'25. 13 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20629 2025-07-29 cs.CV 57%

DAMS:Dual-Branch Adaptive Multiscale Spatiotemporal Framework for Video Anomaly Detection

Dezhi An, Wenqiang Liu, Kefan Wang, Zening Chen, Jun Lu, Shengcai Zhang

机构 * School of Cyberspace Security,Gansu University of Political Science and Law(网络安全学院、政治学科学校)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 13 pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20939 2025-07-29 cs.CV 57%

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li, Teng Wang, Junfu Pu, Yizhuo Li, Lu Qiu, Jin Ma, Lisheng Duan, Xinyu Zuo, Jinwen Luo, Weibo Gu, Zexuan Li, Xiaojing Zhang, Yangyu Tao, Han Hu, Di Wang, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Search Application Department, Tencent CSIG(腾讯CSIG搜索应用部门) Tencent Hunyuan(腾讯文生视频) Big Data Platform Department, Tencent PCG(腾讯PCG大数据平台部门)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://tencentarc.github.io/posts/arc-video-announcement/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20763 2025-07-29 cs.CV 57%

KASportsFormer: Kinematic Anatomy Enhanced Transformer for 3D Human Pose Estimation on Short Sports Scene Video

Zhuoer Yin, Calvin Yeung, Tomohiro Suzuki, Ryota Tanaka, Keisuke Fujii

机构 * Nagoya University(名古屋大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19948 2025-07-29 cs.CV 57%

UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block

Luoxi Jing, Dianxi Shi, Zhe Liu, Songchang Jin, Chunping Qiu, Ziteng Qiao, Yuxian Li, Jianqiang Xia

机构 * School of Computer Science, Peking University(北京大学计算机科学系) Intelligent Game and Decision Lab (IGDL)(智能游戏与决策实验室) College of Computer, National University of Defense Technology(国防科技大学计算机学院) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学系)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IJCAI 2025 (International Joint Conference on Artificial Intelligence)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19863 2025-07-29 cs.MM 57%

Anchoring Trends: Mitigating Social Media Popularity Prediction Drift via Feature Clustering and Expansion

Chia-Ming Lee, Bo-Cheng Qiu, Cheng-Jun Kang, Yi-Hsuan Wu, Jun-Lin Chen, Yu-Fan Lin, Yi-Shiuan Chou, Chih-Chung Hsu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.MM

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01301 2025-07-29 cs.RO cs.AI 57%

Bi-LAT: Bilateral Control-Based Imitation Learning via Natural Language and Action Chunking with Transformers

Takumi Kobayashi, Masato Kobayashi, Thanpimon Buamanee, Yuki Uranishi

机构 * Graduate School of Information Science and Technology, The University of Osaka(信息科学与技术研究生学校,大阪大学) D3 Center, The University of Osaka(大阪大学D3中心) Graduate School of Maritime Sciences, Kobe University(海洋科学研究生学校, Kobe大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19599 2025-07-29 cs.CV 57%

Object-centric Video Question Answering with Visual Grounding and Referring

Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Weidi Xie, Stratis Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) SAI, Shanghai Jiao Tong University(上海交通大学SAI研究所) Xiaohongshu Inc(小红书公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03492 2025-07-29 cs.CV 57%

Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation

Suhwan Cho, Seunghoon Lee, Minhyeok Lee, Jungho Lee, Sangyoun Lee

机构 * GenGenAI Yonsei University(延世大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments ICCVW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21016 2025-07-29 cs.LG q-bio.NC 50%

Predicting Cognition from fMRI:A Comparative Study of Graph, Transformer, and Kernel Models Across Task and Rest Conditions

Jagruti Patel, Mikkel Schöttner, Thomas A. W. Bolton, Patric Hagmann

专题命中 视频多模态 :multimodal(abstract)

Comments Preliminary version; a revised version will be uploaded later

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20866 2025-07-29 physics.comp-ph 50%

Neuromorphic Photonic Processing and Memory with Spiking Resonant Tunnelling Diode Neurons and Neural Networks

Dafydd Owen-Newns, Joshua Robertson, Giovanni Donati, Jose Figueiredo, Edward Wasige, Kathy Ludge, Bruno Romeira, Antonio Hurtado

专题命中 视频多模态 :multi-modal(abstract)

Comments 19 pages, 11 figures, submitted to Advanced Intelligent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19566 2025-07-29 eess.IV 50%

SLENet: A Novel Multiscale CNN-Based Network for Detecting the Rats Estrous Cycle

Qinyang Wang, Hoileong Lee, Xiaodi Pu, Yuanming Lai, Yiming Ma

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏