arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-13 至 2025-10-13 共收录 12 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 3 篇

2510.06509 2025-10-13 cs.CV 85%

From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding

Shih-Yao Lin, Sibendu Paul, Caren Chen

机构 * Amazon Prime Video(亚马逊Prime视频)

专题命中 视频理解 :video-language(title,abstract);video understanding(abstract);long video(abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08818 2025-10-13 cs.CV cs.AI 70%

D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition

Yiyang Huang, Yizhou Wang, Yun Fu

机构 * Northeastern University(东北大学)

专题命中 视频理解 :video understanding(abstract);video-language(abstract);分类 cs.CV

Comments This paper has been accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09230 2025-10-13 cs.CV cs.AI cs.CL cs.LG 57%

Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras

Jindong Hong, Wencheng Zhang, Shiqin Qiao, Jianhai Chen, Jianing Qiu, Chuanyang Zheng, Qian Xu, Yun Ji, Qianyue Wen, Weiwei Sun, Hao Li, Huizhen Li, Huichao Wang, Kai Wu, Meng Li, Yijun He, Lingjie Luo, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) Peking University People’s Hospital(北京大学人民医院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 2 篇

2505.19495 2025-10-13 cs.CV 83%

The Role of Video Generation in Enhancing Data-Limited Action Understanding

Wei Li, Dezhao Luo, Dongbao Yang, Zhenhang Li, Weiping Wang, Yu Zhou

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) VCIP & TMCC & DISSec, College of Computer Science, Nankai University(南开大学计算机学院) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Queen Mary University of London(伦敦大学玛丽女王学院)

专题命中 视频生成 :video generation(title);video diffusion(abstract);text-to-video(abstract);分类 cs.CV

Comments IJCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08799 2025-10-13 cs.CV cs.AI cs.LG 57%

SkipSR: Faster Super Resolution with Token Skipping

Rohan Choudhury, Shanchuan Lin, Jianyi Wang, Hao Chen, Qi Zhao, Feng Cheng, Lu Jiang, Kris Kitani, Laszlo A. Jeni

机构 * Carnegie Mellon University(卡内基梅隆大学) ByteDance Seed(字节跳动种子)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 4 篇

2506.03517 2025-10-13 cs.CV 83%

DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models

Ziyi Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace, Ashkan Mirzaei, Igor Gilitschenski, Sergey Tulyakov, Aliaksandr Siarohin

机构 * Snap Research University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 视频扩散模型 :video diffusion(title,abstract);text-to-video(abstract);分类 cs.CV

Comments NeurIPS 2025 Spotlight. Project page: https://snap-research.github.io/DenseDPO/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21593 2025-10-13 cs.CV cs.AI 79%

Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model

Yang Yang, Siming Zheng, Qirui Yang, Jinwei Chen, Boxi Wu, Xiaofei He, Deng Cai, Bo Li, Peng-Tao Jiang

机构 * Zhejiang University(浙江大学) vivo Mobile Communication Co., Ltd(vivo移动通信有限公司)

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

Comments project page: https://vivocameraresearch.github.io/any2bokeh/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02851 2025-10-13 cs.CV cs.GR 79%

Human-VDM: Learning Single-Image 3D Human Gaussian Splatting from Video Diffusion Models

Zhibin Liu, Haoye Dong, Aviral Chharia, Hefeng Wu

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

Comments 14 Pages, 8 figures, Project page: https://human-vdm.github.io/Human-VDM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09212 2025-10-13 cs.CV 74%

Stable Video Infinity: Infinite-Length Video Generation with Error Recycling

Wuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao, Alexandre Alahi

机构 * EPFL(苏黎世联邦理工学院)

专题命中 视频扩散模型 :video generation(title);分类 cs.CV

Comments Project Page: https://stable-video-infinity.github.io/homepage/

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 长视频与时序推理 2 篇

2510.09182 2025-10-13 cs.CV 57%

Online Video Depth Anything: Temporally-Consistent Depth Prediction with Low Memory Consumption

Johann-Friedrich Feiden, Tim Küchler, Denis Zavadski, Bogdan Savchynskyy, Carsten Rother

机构 * Heidelberg University(海德堡大学)

专题命中 长视频与时序推理 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09553 2025-10-13 cs.CL 50%

Hierarchical Indexing with Knowledge Enrichment for Multilingual Video Corpus Retrieval

Yu Wang, Tianhao Tan, Yifei Wang

机构 * School of Computing and Information, University of Pittsburgh, PA, USA(计算与信息学院,匹兹堡大学) Wuhan University of Technology(武汉理工大学) Hunan University(湖南大学)

专题命中 长视频与时序推理 :long video(abstract)

Comments Accepted to NLPCC 2025 (Springer), to appear November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 视频数据与评测 1 篇

2510.08936 2025-10-13 cs.CV cs.AI 57%

RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos

Zixi Yang, Jiapeng Li, Muxi Diao, Yinuo Jing, Kongming Liang

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏