arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-07-22 至 2025-07-22 共收录 13 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1 篇

2507.15569 2025-07-22 cs.CV 79%

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

Xiaoyi Bao, Chenwei Xie, Hao Tang, Tingyu Weng, Xiaofeng Wang, Yun Zheng, Xingang Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团) Peking University(北京大学) Luoyang Institute for Robot and Intelligent Equipment(洛阳机器人与智能装备研究所)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 2 篇

2507.15824 2025-07-22 cs.CV 79%

Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models

Enes Sanli, Baris Sarper Tezcan, Aykut Erdem, Erkut Erdem

专题命中 视频生成 :video generation(title);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01425 2025-07-22 cs.CV 79%

Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation

Xincheng Shuai, Henghui Ding, Zhenyuan Qin, Hao Luo, Xingjun Ma, Dacheng Tao

机构 * Fudan University(复旦大学) DAMO Academy, Alibaba group(阿里集团 DAMO 院) Hupan Lab(虎派实验室) Nanyang Technological University, Singapore(新加坡南洋理工大学)

专题命中 视频生成 :video generation(title);text-to-video(abstract);分类 cs.CV

Comments ICCV 2025, Project Page: https://henghuiding.github.io/SynFMC/

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 4 篇

2505.00704 2025-07-22 cs.GR cs.CV 79%

Controllable Weather Synthesis and Removal with Video Diffusion Models

Chih-Hao Lin, Zian Wang, Ruofan Liang, Yuxuan Zhang, Sanja Fidler, Shenlong Wang, Zan Gojcic

机构 * NVIDIA University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

Comments International Conference on Computer Vision (ICCV) 2025, Project Website: https://research.nvidia.com/labs/toronto-ai/WeatherWeaver/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08676 2025-07-22 cs.CV cs.GR 70%

FlexiClip: Locality-Preserving Free-Form Character Animation

Anant Khandelwal

机构 * Search Technology Center India, Microsoft, IDC, Bengaluru, India(Bharat)(印度搜索技术中心、微软、IDC、班加罗尔、印度(巴哈拉特))

专题命中 视频扩散模型 :video diffusion(abstract);text-to-video(abstract);分类 cs.CV

Comments 13 pages, 4 figures, 7 tables, Accepted in ICML 2025, https://openreview.net/forum?id=xtxCM4XZ82

Journal ref Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15064 2025-07-22 cs.CV cs.AI 57%

StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation

Shuyuan Tu, Zhen Xing, Xintong Han, Zhi-Qi Cheng, Qi Dai, Chong Luo, Zuxuan Wu, Yu-Gang Jiang

机构 * School of Computer Science, Fudan University(复旦大学计算机科学学院) Microsoft Research Asia(微软亚洲研究院) Tencent Inc.(腾讯公司) University of Washington(华盛顿大学)

专题命中 视频扩散模型 :video diffusion(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2411.17697

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15260 2025-07-22 cs.LG 50%

CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE Solvers

Jiaqi Han, Haotian Ye, Puheng Li, Minkai Xu, James Zou, Stefano Ermon

机构 * Stanford University(斯坦福大学)

专题命中 视频扩散模型 :video diffusion(abstract)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 动作与事件理解 2 篇

2507.15428 2025-07-22 cs.CV cs.AI 79%

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent

Jiaao Li, Kaiyuan Li, Chen Gao, Yong Li, Xinlei Chen

专题命中 动作与事件理解 :video reasoning(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02192 2025-07-22 cs.CV cs.AI 70%

DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization

Wenchuan Wang, Mengqi Huang, Yijing Tu, Zhendong Mao

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 动作与事件理解 :video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 长视频与时序推理 3 篇

2507.15728 2025-07-22 cs.CV 89%

TokensGen: Harnessing Condensed Tokens for Long Video Generation

Wenqi Ouyang, Zeqi Xiao, Danni Yang, Yifan Zhou, Shuai Yang, Lei Yang, Jianlou Si, Xingang Pan

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) SenseTime Research(商汤科技研究院) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)

专题命中 长视频与时序推理 :video generation(title,abstract);long video(title,abstract);video diffusion(abstract);分类 cs.CV

Comments Project page: https://vicky0522.github.io/tokensgen-webpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18423 2025-07-22 cs.CV cs.AI 57%

OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics

Yeon-Ji Song, Jaein Kim, Suhyung Choi, Jin-Hwa Kim, Byoung-Tak Zhang

机构 * Seoul National University(首尔国立大学) SNU AIIS NAVER AI Lab(NAVER AI实验室)

专题命中 长视频与时序推理 :long video(abstract);分类 cs.CV

Comments Accepted at ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05166 2025-07-22 cs.CV 57%

CD-NGP: A Fast Scalable Continual Representation for Dynamic Scenes

Zhenhuan Liu, Shuai Liu, Zhiwei Ning, Jie Yang, Yifan Zuo, Yuming Fang, Wei Liu

机构 * Dept. of Automation, Shanghai Jiao Tong University(自动化系,上海交通大学)

专题命中 长视频与时序推理 :long video(abstract);分类 cs.CV

Comments 18 pages in total

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 视频数据与评测 1 篇

2507.15028 2025-07-22 cs.CV 79%

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

Yuanhan Zhang, Yunice Chew, Yuhao Dong, Aria Leo, Bo Hu, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 视频数据与评测 :video reasoning(title);video understanding(abstract);分类 cs.CV

Comments ICCV 2025; Project page: https://zhangyuanhan-ai.github.io/video-tt/

详情

展开后加载摘要…

URL PDF HTML 收藏