arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6840 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1395 篇

2409.11513 2025-02-03 cs.CV cs.AI 57%

Mamba Fusion: Learning Actions Through Questioning

Zhikang Dong, Apoorva Beedu, Jason Sheinkopf, Irfan Essa

机构 * Georgia Institute of Technology(佐治亚理工学院) Stony Brook University(石溪大学) Google Research(谷歌研究院)

专题命中 视频理解 :video language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07810 2025-01-15 cs.CV 57%

AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation

Sitong Gong, Yunzhi Zhuge, Lu Zhang, Yifan Wang, Pingping Zhang, Lijun Wang, Huchuan Lu

机构 * School of Information and Communication Engineering, Dalian University of Technology(大连理工大学信息与通信工程学院) School of Innovation and Entrepreneurship, Dalian University of Technology(大连理工大学创新创业学院) School of Artificial Intelligence, Dalian University of Technology(大连理工大学人工智能学院)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Multimedia (TMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07087 2025-01-14 cs.CV cs.AI 57%

Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling

Jiebin Yan, Lei Wu, Yuming Fang, Xuelin Liu, Xue Xia, Weide Liu

机构 * School of Information Technology, Jiangxi University of Finance and Economics(江西财经大学信息技术学院) Harvard Medical School, Harvard University(哈佛大学医学院)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22829 2025-01-14 cs.CV 57%

Situational Scene Graph for Structured Human-centric Situation Understanding

Chinthani Sugandhika, Chen Li, Deepu Rajan, Basura Fernando

机构 * College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) Centre for Frontier AI Research, Agency for Science, Technology and Research(新加坡科技研究局前沿人工智能研究中心) Institute of High-Performance Computing, Agency for Science, Technology and Research(新加坡科技研究局高性能计算研究所)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted for WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06761 2025-01-14 cs.CV 57%

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning

Ji Soo Lee, Jongha Kim, Jeehye Na, Jinyoung Park, Hyunwoo J. Kim

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05901 2025-01-14 cs.CV 57%

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Ziheng Wu, Zhenghao Chen, Ruipu Luo, Can Zhang, Yuan Gao, Zhentao He, Xian Wang, Haoran Lin, Minghui Qiu

机构 * ByteDance(字节跳动)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05717 2025-01-13 cs.CV cs.AI q-bio.QM 57%

Zero-shot Shark Tracking and Biometrics from Aerial Imagery

Chinmay K Lalgudi, Mark E Leone, Jaden V Clark, Sergio Madrigal-Mora, Mario Espinoza

机构 * Stanford University(斯坦福大学) Flinders University(弗林德斯大学) Centro de Investigación en Ciencias del Mar y Limnología, Universidad de Costa Rica(哥斯达黎加大学海洋与湖沼科学研究中心)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06433 2025-01-13 cs.CV cs.AI cs.LG 57%

Self-supervised video pretraining yields robust and more human-aligned visual representations

Nikhil Parthasarathy, S. M. Ali Eslami, João Carreira, Olivier J. Hénaff

机构 * Google DeepMind(谷歌DeepMind)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted to 37th Conference on Neural Information Processing Systems (NeurIPS 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00432 2025-01-03 cs.CV cs.LG 57%

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models

Lala Shakti Swarup Ray, Bo Zhou, Sungho Suh, Paul Lukowicz

机构 * RPTU Kaiserslautern-Landau(莱茵兰-普法尔茨州立大学凯撒斯劳滕-兰道分校) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted in IEEE ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00136 2025-01-03 cs.CV cs.AI cs.LG 57%

Detection-Fusion for Knowledge Graph Extraction from Videos

Taniya Das, Louis Mahon, Thomas Lukasiewicz

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 12 pages, To be submitted to a conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17759 2024-12-24 cs.AI cs.CV cs.LG 57%

Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy

Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Bhargava Kumar, Amit Agarwal, Ishan Banerjee, Srikant Panda, Tejaswini Kumar

机构 * University of Washington(华盛顿大学) New York University(纽约大学) Columbia University(哥伦比亚大学) Liverpool John Moores University(利物浦约翰摩尔大学) Chennai Mathematical Institute(金奈数学研究所) Birla Institute of Technology(比拉理工学院)

专题命中 视频理解 :text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16946 2024-12-24 cs.CV 57%

Video Domain Incremental Learning for Human Action Recognition in Home Environments

Yuanda Hu, Xing Liu, Meiying Li, Yate Ge, Xiaohua Sun, Weiwei Guo

机构 * College of Design and Innovation, Tongji University(同济大学设计与创新学院) SUSTech School of Design, Southern University of Science and Technology(南方科技大学设计学院)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16117 2024-12-23 cs.CV 57%

PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Xiaohu Huang, Hao Zhou, Kai Han

机构 * The University of Hong Kong(香港大学) Baidu Inc.(百度公司)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Efficient Video Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14006 2024-12-19 cs.CV 57%

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

Cong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng, Yong Liu, Zheng Zhao, Yujiu Yang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Meituan Inc.(美团公司)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12675 2024-12-18 cs.CV 57%

ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries

Wangyu Xue, Chen Qian, Jiayi Wu, Yang Zhou, Wentao Liu, Ju Ren, Siming Fan, Yaoxue Zhang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11228 2024-12-17 cs.CV cs.AI cs.LG 57%

Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition

Yulin Wang, Haoji Zhang, Yang Yue, Shiji Song, Chao Deng, Junlan Feng, Gao Huang

机构 * Tsinghua University(清华大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) China Mobile Research Institute(中国移动研究院)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted by IEEE TPAMI. Journal version of arXiv:2105.03245 (AdaFocusV1, ICCV 2021 Oral), arXiv:2112.14238 (AdaFocusV2, CVPR 2022), and arXiv:2209.13465 (AdaFocusV3, ECCV 2022). Code and pre-trained models: https://github.com/LeapLabTHU/Uni-AdaFocus

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12519 2024-12-10 cs.CV 57%

Dynamics Based Neural Encoding with Inter-Intra Region Connectivity

Mai Gamal, Mohamed Rashad, Eman Ehab, Seif Eldawlatly, Mennatullah Siam

机构 * German University in Cairo(开罗德国大学) Ain Shams University(艾因夏姆斯大学) Nile University(尼罗河大学) American University in Cairo(开罗美国大学) University of British Columbia(不列颠哥伦比亚大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Title change

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19644 2024-11-28 cs.CV cs.AI cs.LG 57%

EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos

Ryo Fujii, Masashi Hatano, Hideo Saito, Hiroki Kajita

机构 * Keio University(庆应义塾大学) Keio University School of Medicine(庆应义塾大学医学院)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Early accepted by MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14205 2024-11-22 cs.CV cs.AI 57%

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body

Zeqing Wang, Qingyang Ma, Wentao Wan, Haojie Li, Keze Wang, Yonghong Tian

专题命中 视频理解 :text-to-video(abstract);分类 cs.CV

Comments 16 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13683 2024-11-22 cs.CV 57%

Extending Video Masked Autoencoders to 128 frames

Nitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal, Chaitra Hegde, Eirikur Agustsson, Sagar M. Waghmare, Mikhail Sirotenko, Ming-Hsuan Yang, Tobias Weyand, Boqing Gong, Leonid Sigal

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 10.5 pages of main paper, 25 pages total, 4 figures and 10 tables. To appear in NeurIPS'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11066 2024-11-19 cs.CV 57%

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Tingyu Qu, Mingxiao Li, Tinne Tuytelaars, Marie-Francine Moens

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13800 2024-11-18 cs.CV cs.AI 57%

Dense Connector for MLLMs

Huanjin Yao, Wenhao Wu, Taojiannan Yang, YuXin Song, Mengxi Zhang, Haocheng Feng, Yifan Sun, Zhiheng Li, Wanli Ouyang, Jingdong Wang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 27 pages, NeurIPS 2024

Journal ref NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19006 2024-11-14 cs.CV 57%

Snakes and Ladders: Two Steps Up for VideoMamba

Hui Lu, Albert Ali Salah, Ronald Poppe

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments New updated experiment results

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05636 2024-11-11 cs.CV cs.LG 57%

Video RWKV:Video Action Recognition Based RWKV

Zhuowen Yin, Chengru Li, Xingbo Dong

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01335 2024-11-04 cs.CV cs.AI 57%

BehAVE: Behaviour Alignment of Video Game Encodings

Nemanja Rašajski, Chintan Trivedi, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03326 2024-10-29 cs.CV cs.AI cs.CL 57%

LLaVA-OneVision: Easy Visual Task Transfer

Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, Chunyuan Li

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Project Homepage: https://llava-vl.github.io/blog/2024-08-05-llava-onevision/

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.17327 2024-10-24 cs.CV 57%

Telling Stories for Common Sense Zero-Shot Action Recognition

Shreyank N Gowda, Laura Sevilla-Lara

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted in ACCV 2024!

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10441 2024-10-17 cs.CV cs.AI 57%

Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs

Kai Han, Jianyuan Guo, Yehui Tang, Wei He, Enhua Wu, Yunhe Wang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09875 2024-10-15 cs.CV cs.IR 57%

ViFi-ReID: A Two-Stream Vision-WiFi Multimodal Approach for Person Re-identification

Chen Mao, Chong Tan, Jingqi Hu, Min Zheng

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07405 2024-10-11 cs.CV cs.AI 57%

Exploring Efficient Foundational Multi-modal Models for Video Summarization

Karan Samel, Apoorva Beedu, Nitish Sontakke, Irfan Essa

专题命中 视频理解 :video language model(abstract);分类 cs.CV

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏