arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-11-06 至 2025-11-06 共收录 7 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1 篇

2503.18422 2025-11-06 cs.CV 83%

Breaking the Encoder Barrier for Seamless Video-Language Understanding

Handong Li, Yiyuan Zhang, Longteng Guo, Xiangyu Yue, Jing Liu

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MMLab, CUHK(CUHK MMLab) Institute of Automation, Chinese Academy of Science(中国科学院自动化研究所) Shanghai AI Lab(上海人工智能实验室)

专题命中 视频理解 :video-language(title,abstract);video understanding(abstract);分类 cs.CV

Comments 12 pages

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 23167-23176

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 2 篇

2506.09045 2025-11-06 cs.CV 79%

MagCache: Fast Video Generation with Magnitude-Aware Cache

Zehong Ma, Longhui Wei, Feng Wang, Shiliang Zhang, Qi Tian

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) Huawei Inc.(华为公司)

专题命中 视频生成 :video generation(title);video diffusion(abstract);分类 cs.CV

Comments Project Page: https://zehong-ma.github.io/MagCache Accepted by NeurIPS 2025

Journal ref In Proceedings of NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14799 2025-11-06 cs.CV cs.AI 57%

A Survey on Text-Driven 360-Degree Panorama Generation

Hai Wang, Xiaoyu Xiang, Weihao Xia, Jing-Hao Xue

机构 * Department of Statistical Science, University College London(统计科学系,伦敦大学学院) Core AI team at Meta Reality Labs(Meta Reality Labs人工智能核心团队)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted by IEEE TCSVT, Code: https://github.com/littlewhitesea/Text-Driven-Pano-Gen

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 2 篇

2511.03272 2025-11-06 cs.CV 87%

Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising

Shuangquan Lyu, Steven Mao, Yue Ma

专题命中 视频扩散模型 :long video(title,abstract);video generation(abstract);video diffusion(abstract);text-to-video(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10687 2025-11-06 cs.CV 74%

Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation

Hao Zhang, Chun-Han Yao, Simon Donné, Narendra Ahuja, Varun Jampani

机构 * Stability AI University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视频扩散模型 :video generation(title);分类 cs.CV

Comments Page: https://stablepartdiffusion4d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频问答 1 篇

2510.08668 2025-11-06 cs.CV 57%

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Songtao Jiang, Yuan Wang, Sibo Song, Tianxiang Hu, Chenyi Zhou, Bin Pu, Yan Zhang, Zhibo Yang, Yang Feng, Joey Tianyi Zhou, Jin Hao, Zijian Chen, Ruijia Wu, Tao Tang, Junhui Lv, Hongxia Xu, Hongwei Wang, Jun Xiao, Bin Feng, Fudong Zhu, Kenli Li, Weidi Xie, Jimeng Sun, Jian Wu, Zuozhu Liu

机构 * College of Computer Science and Technology, Zhejiang University-University of Illinois Urbana-Champaign Institute(浙江大学计算机科学与技术学院) Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine(浙江大学口腔医院) Alibaba Inc(阿里巴巴集团) College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院) Angelalign Technology Inc.(Angelalign技术有限公司) CFAR & IHPC, Agency for Science, Technology and Research(CFAR与IHPC,新加坡科技研究局) Department of Orthodontics, Shanghai Ninth People’s Hospital, College of Stomatology, Shanghai Jiao Tong University(上海第九人民医院正畸科,上海交通大学口腔医学院)

专题命中 视频问答 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 视频数据与评测 1 篇

2511.03178 2025-11-06 cs.CV 57%

SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention

Shreyas C. Dhake, Jiayuan Huang, Runlong He, Danyal Z. Khan, Evangelos B. Mazomenos, Sophia Bano, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarak I. Hoque

机构 * UCL Hawkes Institute(UCL哈维斯研究所) University College London(伦敦大学学院) Dept of Medical Physics & Biomedical Engineering(医学物理与生物医学工程系) UCL(伦敦大学学院) Dept of Computer Science(计算机科学系) National Hospital for Neurology and Neurosurgery(神经病学与神经外科国家医院) Division of Informatics, Imaging and Data Science(信息学、成像与数据科学 division)

专题命中 视频数据与评测 :video language model(abstract);分类 cs.CV

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏