arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-06 至 2025-10-06 共收录 11 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2510.02778 2025-10-06 cs.CV 79%

AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding

Xian Zhang, Zexi Wu, Zinuo Li, Hongming Xu, Luqi Gong, Farid Boussaid, Naoufel Werghi, Mohammed Bennamoun

机构 * The University of Western Australia(西澳大学) Dalian University of Technology(大连理工大学) Khalifa University(卡利夫大学) Zhejiang Lab(浙江实验室)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08665 2025-10-06 cs.CV 74%

SkillFormer: Unified Multi-View Video Understanding for Proficiency Estimation

Edoardo Bianchi, Antonio Liotta

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments Accepted at the 2025 18th International Conference on Machine Vision. Project page at https://edowhite.github.io/SkillFormer

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 3 篇

2510.03049 2025-10-06 cs.CV cs.AI 85%

When and Where do Events Switch in Multi-Event Video Generation?

Ruotong Liao, Guowen Huang, Qing Cheng, Thomas Seidl, Daniel Cremers, Volker Tresp

机构 * Ludwig-Maxilians-University of Munich(慕尼黑路德维希-马克西米利安大学) Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 视频生成 :video generation(title,abstract);text-to-video(abstract);long video(abstract);分类 cs.CV

Comments Work in Progress. Accepted to ICCV2025 @ LongVid-Foundations

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01784 2025-10-06 cs.CV cs.AI 79%

Pack and Force Your Memory: Long-form and Consistent Video Generation

Xiaofei Wu, Guozhen Zhang, Zhiyong Xu, Yuan Zhou, Qinglin Lu, Xuming He

机构 * ShanghaiTech University(上海科技大学) Tencent Hunyuan(腾讯文言) Nanjing University(南京大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18406 2025-10-06 cs.CV cs.AI cs.CL 57%

RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives

Jaehong Yoon, Shoubin Yu, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学)

专题命中 视频生成 :video diffusion(abstract);分类 cs.CV

Comments EMNLP 2025 main; The first two authors contribute equally. Project Page: https://raccoon-mllm-gen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 2 篇

2510.01669 2025-10-06 cs.CV 79%

UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction

Jin Cao, Hongrui Wu, Ziyong Feng, Hujun Bao, Xiaowei Zhou, Sida Peng

机构 * State Key Lab of CAD&CG, Zhejiang University(CAD与CG国家重点实验室,浙江大学) Tongji University(同济大学) DeepGlint

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

Comments page: https://jin-cao-tma.github.io/UniVerse.github.io/ code: https://github.com/zju3dv/UniVerse

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09650 2025-10-06 cs.CV cs.LG cs.MM cs.RO eess.IV 67%

HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios

Kunyu Peng, Junchao Huang, Xiangsheng Huang, Di Wen, Junwei Zheng, Yufan Chen, Kailun Yang, Jiamin Wu, Chongqing Hao, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Hunan University(湖南大学) Shanghai AI Lab(上海人工智能实验室) HEBUST

专题命中 视频扩散模型 :video understanding(abstract);分类 cs.CV、eess.IV、cs.MM

Comments Accepted to NeurIPS 2025. The dataset and code are available at https://github.com/KPeng9510/HopaDIFF

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 长视频与时序推理 3 篇

2510.02617 2025-10-06 cs.CV 74%

Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation

Beijia Lu, Ziyi Chen, Jing Xiao, Jun-Yan Zhu

机构 * Carnegie Mellon University(卡内基梅隆大学) PAII Inc.(PAII公司)

专题命中 长视频与时序推理 :video generation(title);分类 cs.CV

Comments Project Page: https://beijia11.github.io/IASA

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02790 2025-10-06 cs.CV cs.CL 74%

From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding

Xiangfeng Wang, Xiao Li, Yadong Wei, Xueyu Song, Yang Song, Xiaoqiang Xia, Fangrui Zeng, Zaiyi Chen, Liu Liu, Gu Xu, Tong Xu

机构 * University of Science and Technology of China(中国科学技术大学) ByteDance China(字节跳动中国)

专题命中 长视频与时序推理 :long video(title);分类 cs.CV

Comments Accepted by EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03198 2025-10-06 cs.CV 57%

Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft

Junchao Huang, Xinting Hu, Boyao Han, Shaoshuai Shi, Zhuotao Tian, Tianyu He, Li Jiang

专题命中 长视频与时序推理 :video diffusion(abstract);分类 cs.CV

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 视频数据与评测 1 篇

2510.02571 2025-10-06 cs.CV cs.AI cs.CL 70%

How Confident are Video Models? Empowering Video Models to Express their Uncertainty

Zhiting Mei, Ola Shorinwa, Anirudha Majumdar

机构 * Princeton University(普林斯顿大学)

专题命中 视频数据与评测 :video generation(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏