arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2508.15717 2025-08-22 cs.CV cs.AI 62%

StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

Yanlai Yang, Zhuokai Zhao, Satya Narayan Shukla, Aashu Singh, Shlok Kumar Mishra, Lizhu Zhang, Mengye Ren

机构 * Meta AI New York University(纽约大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14941 2025-08-22 cs.MM cs.CL 62%

Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs

Yi-Chun Chen

机构 * Yale University(耶鲁大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.MM

Comments 12 pages, 4 figures, 2 tables. Extends our earlier framework on hierarchical narrative graphs with a semantic normalization module

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01712 2025-08-18 cs.CV cs.AI 62%

HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection

Han Wang, Zhuoran Wang, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08035 2025-08-12 cs.CV cs.AI 62%

LVBench: An Extreme Long Video Understanding Benchmark

Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Xiaotao Gu, Shiyu Huang, Bin Xu, Yuxiao Dong, Ming Ding, Jie Tang

机构 * Zhipu AI(智谱AI) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15384 2025-08-11 cs.CV cs.AI 62%

MetaOcc: Spatio-Temporal Fusion of Surround-View 4D Radar and Camera for 3D Occupancy Prediction with Dual Training Strategies

Long Yang, Lianqing Zheng, Wenjin Ai, Minghao Liu, Sen Li, Qunshu Lin, Shengyu Yan, Jie Bai, Zhixiong Ma, Tao Huang, Xichan Zhu

机构 * School of Automotive Studies, Tongji University(同济大学汽车学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Automobile, Chang'an University(长安大学汽车学院) School of Information and Electrical Engineering, Hangzhou City University(杭州城市学院信息与电气工程学院) College of Science and Engineering, James Cook University(詹姆斯库克大学科学与工程学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03009 2025-08-06 cs.CV cs.AI 62%

Enhancing Long Video Question Answering with Scene-Localized Frame Grouping

Xuyi Yang, Wenhao Zhang, Hongbo Jin, Lin Liu, Hongbo Xu, Yongwei Nie, Fei Yu, Fei Ma

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18688 2025-08-06 cs.CV cs.AI 62%

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

Faraz Waseem, Muhammad Shahzad

机构 * University Of Reading(阅读大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 35 pages, 18 figures, Manuscript submitted to ACM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01781 2025-08-05 cs.CL cs.AI 62%

A comprehensive taxonomy of hallucinations in Large Language Models

Manuel Cossio

机构 * Universitat de Barcelona(巴塞罗那大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 55 pages, 16 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00085 2025-08-04 cs.CV cs.AI 62%

Punching Bag vs. Punching Person: Motion Transferability in Videos

Raiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran, Yogesh Rawat

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) Center for Vision Technology, SRI International(视觉技术中心,SRI国际)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02713 2025-08-04 cs.CV cs.CL 62%

LLaVA-Video: Video Instruction Tuning With Synthetic Data

Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, Chunyuan Li

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) BUPT(北京邮电大学) ByteDance(字节跳动)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Project page: https://llava-vl.github.io/blog/2024-09-30-llava-video/; Accepted at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.01430 2025-08-01 cs.CV cs.AI cs.LG 62%

Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots

Dong Lao, Zhengyang Hu, Francesco Locatello, Yanchao Yang, Stefano Soatto

机构 * UCLA(加州大学洛杉矶分校) HKU(香港大学) ISTA(因斯布鲁克大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20947 2025-08-01 cs.CV cs.MM 62%

Hierarchical Sub-action Tree for Continuous Sign Language Recognition

Dejie Yang, Zhu Xu, Xinjie Gao, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University, Beijing, China(王轩计算机技术研究所,北京大学,北京,中国) State Key Laboratory of General Artificial Intelligence, Peking Universitys, Beijing, China(通用人工智能国家重点实验室,北京大学,北京,中国)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Journal ref ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17402 2025-07-29 cs.CV cs.IR cs.MM 62%

HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

Jun Li, Jinpeng Wang, Chaolei Tan, Niu Lian, Long Chen, Yaowei Wang, Min Zhang, Shu-Tao Xia, Bin Chen

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Research Center of Artificial Intelligence, Peng Cheng Laboratory(鹏城实验室人工智能研究中心) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by ICCV'25. 13 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17958 2025-07-28 cs.LG cs.AI cs.CV 62%

VIBE: Video-Input Brain Encoder for fMRI Response Modeling

Daniel Carlström Schad, Shrey Dixit, Janis Keck, Viktor Studenyak, Aleksandr Shpilevoi, Andrej Bicanski

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16151 2025-07-23 cs.CV cs.AI 62%

SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities

Yasser Ashraf, Ahmed Sharshar, Velibor Bojkovic, Bin Gu

机构 * Department of Machine Learning(机器学习系) Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13820 2025-07-21 cs.CV cs.AI 62%

Team of One: Cracking Complex Video QA with Model Synergy

Jun Xie, Zhaoran Zhao, Xiongjun Guan, Yingjian Zhu, Hongzhu Yi, Xinming Wang, Feng Chen, Zhepeng Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12816 2025-07-18 cs.CV cs.AI 62%

FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering

Ju-Young Oh, Ho-Joong Kim, Seong-Whan Lee

机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments SMC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08411 2025-07-15 cs.LG cs.AI cs.CV stat.AP 62%

BiDepth: A Bidirectional-Depth Neural Network for Spatio-Temporal Prediction

Sina Ehsani, Fenglian Pan, Qingpei Hu, Jian Liu

机构 * University of Arizona(亚利桑那大学) Chinese Academy of Sciences(中国科学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 6 figures. Submitted to ACM TKDD

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06523 2025-07-10 cs.CV cs.CL cs.GR 62%

FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation

Liqiang Jing, Viet Lai, Seunghyun Yoon, Trung Bui, Xinya Du

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04976 2025-07-08 cs.CV cs.CL 62%

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

Eunseop Yoon, Hee Suk Yoon, Mark A. Hasegawa-Johnson, Chang D. Yoo

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) University of Illinois at Urbana-Champaign (UIUC)(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13860 2025-07-08 cs.CV cs.AI 62%

Domain Adaptation of VLM for Soccer Video Understanding

Tiancheng Jiang, Henry Wang, Md Sirajus Salekin, Parmida Atighehchian, Shinan Zhang

机构 * Massachusetts Institute of Technology(麻省理工学院) Amazon Web Services(亚马逊网络服务)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 5 figures, accepted to the 11th IEEE International Workshop on Computer Vision in Sports (CVSports) at CVPR 2025; supplementary appendix included

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00950 2025-07-02 cs.CV cs.LG cs.MM 62%

MVP: Winning Solution to SMP Challenge 2025 Video Track

Liliang Ye, Yunyao Zhang, Yafeng Wu, Yi-Ping Phoebe Chen, Junqing Yu, Wei Yang, Zikai Song

机构 * Huazhong University of Science and Technology(华中科技大学) La Trobe University(拉特罗布大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09105 2025-07-02 cs.CV cs.AI 62%

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

Chenglin Li, Qianglong Chen, Zhi Li, Feng Tao, Yin Zhang

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21272 2025-06-30 cs.GR cs.CV cs.MM 62%

FairyGen: Storied Cartoon Video from a Single Child-Drawn Character

Jiayi Zheng, Xiaodong Cun

机构 * GVC Lab, Great Bay University(Great Bay大学GVC实验室)

专题命中 视频多模态 :MLLM(abstract);分类 cs.CV、cs.MM

Comments Project Page: https://jayleejia.github.io/FairyGen/ ; Code: https://github.com/GVCLab/FairyGen

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18071 2025-06-30 cs.CV cs.AI 62%

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering

Jisheng Dang, Huilin Song, Junbin Xiao, Bimei Wang, Han Peng, Haoxuan Li, Xun Yang, Meng Wang, Tat-Seng Chua

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21080 2025-06-27 cs.CV cs.AI cs.LG 62%

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

Sanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan, Calvin Murdock, Ishwarya Ananthabhotla, Yijun Qian, Vamsi Krishna Ithapu, Dinesh Manocha, Ruohan Gao

机构 * University of Maryland, College Park(马里兰大学学院市分校) Meta Reality Labs(Meta现实实验室) Worcester Polytechnic Institute(沃斯特理工学院) University of Toronto(多伦多大学) FAIR, Meta AI(Meta AI)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20342 2025-06-26 cs.CV cs.AI cs.LG 62%

Feature Hallucination for Self-supervised Action Recognition

Lei Wang, Piotr Koniusz

机构 * Griffith University(格里菲斯大学) Data61/CSIRO(Data61/澳大利亚联邦科学与工业研究组织) University of New South Wales(新南威尔士大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted for publication in International Journal of Computer Vision (IJCV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00526 2025-06-24 cs.CV cs.AI cs.GR 62%

Human Action CLIPs: Detecting AI-generated Human Motion

Matyas Bohacek, Hany Farid

机构 * Google(谷歌) Stanford University(斯坦福大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Journal ref Workshop on Deepfake Detection, Localization and Interpretability @ IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15835 2025-06-23 eess.IV cs.AI cs.CV 62%

MoNetV2: Enhanced Motion Network for Freehand 3D Ultrasound Reconstruction

Mingyuan Luo, Xin Yang, Zhongnuo Yan, Yan Cao, Yuanji Zhang, Xindi Hu, Jin Wang, Haoxuan Ding, Wei Han, Litao Sun, Dong Ni

机构 * National-Regional Key Technology Engineering Laboratory for Medical Ultrasound, School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University, Shenzhen, Guangdong, China(国家级医学超声关键技术研发实验室、生物医学工程学院、深圳大学医学院、深圳大学、深圳、广东、中国) Medical UltraSound Image Computing (MUSIC) Lab, Shenzhen University, Shenzhen, Guangdong, China(医学超声图像计算(MUSIC)实验室、深圳大学、深圳、广东、中国) Shenzhen RayShape Medical Technology Inc.(深圳RayShape医疗科技有限公司) Cancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People’s Hospital, Affiliated People’s Hospital of Hangzhou Medical College, Hangzhou, Zhejiang, China(肿瘤中心、超声医学科、浙江省人民医院、杭州医学院附属人民医院、杭州、浙江、中国) Department of Health Management Center, Qilu Hospital, Cheeloo College of Medicine, Shandong University, Jinan, Shandong, China(健康管理中心、齐鲁医院、山东大学齐鲁医学院、济南、山东、中国)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14144 2025-06-18 cs.CV cs.AI 62%

SceneAware: Scene-Constrained Pedestrian Trajectory Prediction with LLM-Guided Walkability

Juho Bai, Inwook Shim

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏