arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4735 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4735 篇

2505.00681 2025-05-02 cs.LG cs.CV 57%

MINERVA: Evaluating Complex Video Reasoning

Arsha Nagrani, Sachit Menon, Ahmet Iscen, Shyamal Buch, Ramin Mehran, Nilpa Jha, Anja Hauth, Yukun Zhu, Carl Vondrick, Mikhail Sirotenko, Cordelia Schmid, Tobias Weyand

机构 * Google DeepMind(谷歌DeepMind) Columbia University(哥伦比亚大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21853 2025-05-01 cs.CV 57%

A Survey of Interactive Generative Video

Jiwen Yu, Yiran Qin, Haoxuan Che, Quande Liu, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, Hao Chen, Xihui Liu

机构 * The University of Hong Kong, Pok Fu Lam, Hong Kong(香港大学) Kuaishou Technology, Shenzhen, China(快手科技) The Hong Kong University of Science and Technology (HKUST), Hong Kong(香港科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20091 2025-05-01 cs.CV cs.MA 57%

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

Noriyuki Kugo, Xiang Li, Zixin Li, Ashish Gupta, Arpandeep Khatua, Nidhish Jain, Chaitanya Patel, Yuta Kyuragi, Yasunori Ishii, Masamoto Tanabiki, Kazuki Kozuka, Ehsan Adeli

机构 * Panasonic Connect Co., Ltd.(松下电器(中国)有限公司) Stanford University(斯坦福大学) Panasonic R&D Company of America(松下美国研发公司) Panasonic Holdings Corporation(松下控股公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19828 2025-04-29 cs.CV 57%

HOIGaze: Gaze Estimation During Hand-Object Interactions in Extended Reality Exploiting Eye-Hand-Head Coordination

Zhiming Hu, Daniel Haeufle, Syn Schmitt, Andreas Bulling

机构 * University of Stuttgart Germany(斯图加特大学) University of Tuebingen Germany(图宾根大学) The Center for Bionic Intelligence Tuebingen Stuttgart Germany(图宾根-斯图加特生物智能中心)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at SIGGRAPH 2025, link: https://zhiminghu.net/hu25_hoigaze.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18478 2025-04-29 cs.CV 57%

Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Xiangrui Liu, Yan Shu, Zheng Liu, Ao Li, Yang Tian, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) University of Trento(特伦多大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10195 2025-04-29 cs.CV cs.NE q-bio.NC 57%

ST-FlowNet: An Efficient Spiking Neural Network for Event-Based Optical Flow Estimation

Hongze Sun, Jun Wang, Wuque Cai, Duo Chen, Qianqian Liao, Jiayi He, Yan Cui, Dezhong Yao, Daqing Guo

机构 * Clinical Hospital of Chengdu Brain Science Institute(成都脑科学研究院临床医院) MOE Key Lab for NeuroInformation(教育部神经信息关键实验室) China-Cuba Belt and Road Joint Laboratory on Neurotechnology and Brain-Apparatus Communication(中 cuba 带路神经技术与脑-装置通信联合实验室) School of Life Science and Technology, University of Electronic Science and Technology of China(电子科技大学生命科学与技术学院) School of Artificial Intelligence, Chongqing University of Education(重庆教育学院人工智能学院) Sichuan Academy of Medical Sciences and Sichuan Provincial People’s Hospital(四川省医学科学院和四川省人民医院) Research Unit of NeuroInformation (2019RU035), Chinese Academy of Medical Sciences(神经信息研究单位(2019RU035),中国医学科学院)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 13 pages, 6 figures, 6 tables; This work has been submitted to Neural Networks for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18968 2025-04-29 cs.CV cs.LG 57%

Perception of Visual Content: Differences Between Humans and Foundation Models

Nardiena A. Pratama, Shaoyang Fan, Gianluca Demartini

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 12 pages (including references), 5 figures, 5 tables, and a paper/ethics checklist. Camera-Ready Copy for ICWSM 2025. This version uses the same results as the previously posted revise-and-resubmit version. Changes are mostly formatting adjustments

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14856 2025-04-29 cs.CV cs.HC cs.LG 57%

Accessible, At-Home Detection of Parkinson's Disease via Multi-task Video Analysis

Md Saiful Islam, Tariq Adnan, Jan Freyberg, Sangwu Lee, Abdelrahman Abdelkader, Meghan Pawlik, Cathe Schwartz, Karen Jaffe, Ruth B. Schneider, E Ray Dorsey, Ehsan Hoque

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08585 2025-04-25 cs.CV 57%

HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding

Shehreen Azad, Vibhav Vineet, Yogesh Singh Rawat

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) Microsoft Research(微软研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16471 2025-04-24 cs.CV 57%

RGB-D Video Object Segmentation via Enhanced Multi-store Feature Memory

Boyue Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16304 2025-04-24 cs.CV 57%

MetaHarm: Harmful YouTube Video Dataset Annotated by Domain Experts, GPT-4-Turbo, and Crowdworkers

Wonjeong Jo, Magdalena Wojcieszak

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13915 2025-04-22 cs.CV 57%

Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding

Dibyadip Chatterjee, Edoardo Remelli, Yale Song, Bugra Tekin, Abhay Mittal, Bharat Bhatnagar, Necati Cihan Camgöz, Shreyas Hampali, Eric Sauser, Shugao Ma, Angela Yao, Fadime Sener

机构 * Meta Reality Labs(Meta现实实验室) FAIR, Meta(Meta人工智能研究机构) National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 13 pages, 5 figures; https://dibschat.github.io/ProVideLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00584 2025-04-18 cs.CV cs.LG 57%

Online Video Understanding: OVBench and VideoChat-Online

Zhenpeng Huang, Xinhao Li, Jiaqi Li, Jing Wang, Xiangyu Zeng, Cheng Liang, Tao Wu, Xi Chen, Liang Li, Limin Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR 2025 Camera Ready Version. Project Page: https://videochat-online.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12264 2025-04-17 cs.CV 57%

Towards Learning to Complete Anything in Lidar

Ayca Takmaz, Cristiano Saltori, Neehar Peri, Tim Meinhardt, Riccardo de Lutio, Laura Leal-Taixé, Aljoša Ošep

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05810 2025-04-16 cs.CV 57%

PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning

Xinpeng Ding, Kui Zhang, Jianhua Han, Lanqing Hong, Hang Xu, Xiaomeng Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09641 2025-04-15 cs.CV 57%

TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Xingjian Zhang, Siwei Wen, Wenjun Wu, Lei Huang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20384 2025-04-15 cs.RO cs.AI 57%

MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Rongyu Zhang, Menghang Dong, Yuan Zhang, Liang Heng, Xiaowei Chi, Gaole Dai, Li Du, Yuan Du, Shanghang Zhang

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11093 2025-04-15 cs.CV 57%

Text-Promptable Propagation for Referring Medical Image Sequence Segmentation

Runtian Yuan, Mohan Chen, Jilan Xu, Ling Zhou, Qingqiu Li, Yuejie Zhang, Rui Feng, Tao Zhang, Shang Gao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08714 2025-04-15 cs.CV 57%

Selective Query-guided Debiasing for Video Corpus Moment Retrieval

Sunjae Yoon, Ji Woo Hong, Eunseop Yoon, Dahyun Kim, Junyeong Kim, Hee Suk Yoon, Chang D. Yoo

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 16 pages, 6 figures, Accepted in ECCV 2022

Journal ref In European Conference on Computer Vision (pp. 185-200). Springer, Cham (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14174 2025-04-14 cs.CV 57%

Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding

Jingjing Hu, Dan Guo, Kun Li, Zhan Si, Xun Yang, Xiaojun Chang, Meng Wang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06581 2025-04-08 cs.NI cs.CV cs.LG 57%

A Survey on Video Analytics in Cloud-Edge-Terminal Collaborative Systems

Linxiao Gong, Hao Yang, Gaoyun Fang, Bobo Ju, Juncen Guo, Xiaoguang Zhu, Xiping Hu, Yan Wang, Peng Sun, Azzedine Boukerche

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07298 2025-04-02 cs.CV 57%

ALLVB: All-in-One Long Video Understanding Benchmark

Xichen Tan, Yuanjing Luo, Yunfan Ye, Fang Liu, Zhiping Cai

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18565 2025-04-02 cs.CR cs.AI cs.HC 57%

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos

Lehao Lin, Ke Wang, Maha Abdallah, Wei Cai

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments 22 pages, 15 figures; references added, typos corrected, new keyword "guided" added, new experimental data and related results updated; new keyword "Generative AI" added for clarity

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23881 2025-04-01 cs.CV 57%

ExScene: Free-View 3D Scene Reconstruction with Gaussian Splatting from a Single Image

Tianyi Gong, Boyan Li, Yifei Zhong, Fangxin Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23050 2025-04-01 cs.LG cs.CV 57%

Prediction of 30-day hospital readmission with clinical notes and EHR information

Tiago Almeida, Plinio Moreno, Catarina Barata

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08646 2025-04-01 cs.CV 57%

StreamChat: Chatting with Streaming Video

Jihao Liu, Zhiding Yu, Shiyi Lan, Shihao Wang, Rongyao Fang, Jan Kautz, Hongsheng Li, Jose M. Alvare

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18042 2025-04-01 cs.CV 57%

HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation

Trong-Thuan Nguyen, Pha Nguyen, Jackson Cothren, Alper Yilmaz, Khoa Luu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20776 2025-03-31 cs.CV 57%

Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields

Shijie Zhou, Hui Ren, Yijia Weng, Shuwang Zhang, Zhen Wang, Dejia Xu, Zhiwen Fan, Suya You, Zhangyang Wang, Leonidas Guibas, Achuta Kadambi

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20540 2025-03-27 cs.CV 57%

Beyond Intermediate States: Explaining Visual Redundancy through Language

Dingchen Yang, Bowen Cao, Anran Zhang, Weibo Gu, Winston Hu, Guang Chen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20315 2025-03-27 cs.CV 57%

SpikeDerain: Unveiling Clear Videos from Rainy Sequences Using Color Spike Streams

Hanwen Liang, Xian Zhong, Wenxuan Liu, Yajing Zheng, Wenxin Huang, Zhaofei Yu, Tiejun Huang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏