arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-08-13 至 2025-08-13 共收录 5 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1 篇

2508.08989 2025-08-13 cs.CV 83%

KFFocus: Highlighting Keyframes for Enhanced Video Understanding

Ming Nie, Chunwei Wang, Hang Xu, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院)

专题命中 视频理解 :video understanding(title,abstract);long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 1 篇

2508.08588 2025-08-13 cs.CV eess.IV 86%

RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space

Jingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao, Lei Sun, Yichen Qian, Weihua Chen, Fan Wang

机构 * DAMO Academy, Alibaba Group(阿里达摩院) Hupan Lab(华盘实验室) INSAIT Zhejiang University(浙江大学)

专题命中 视频生成 :video generation(title,abstract);video diffusion(abstract);text-to-video(abstract);分类 cs.CV、eess.IV

Comments Project page: https://jingyunliang.github.io/RealisMotion

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 1 篇

2508.08978 2025-08-13 cs.CV 79%

TaoCache: Structure-Maintained Video Generation Acceleration

Zhentao Fan, Zongzuo Wang, Weiwei Zhang

机构 * Huawei Inc.(华为公司)

专题命中 视频扩散模型 :video generation(title);video diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 动作与事件理解 1 篇

2411.13552 2025-08-13 cs.CV 57%

REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion Latents

Rui Tian, Qi Dai, Jianmin Bao, Kai Qiu, Yifan Yang, Chong Luo, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Microsoft Research(微软研究院)

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

Comments Accepted to ICCV2025. Code available at https://github.com/microsoft/Reducio-VAE

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 长视频与时序推理 1 篇

2506.03141 2025-08-13 cs.CV 88%

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Jiwen Yu, Jianhong Bai, Yiran Qin, Quande Liu, Xintao Wang, Pengfei Wan, Di Zhang, Xihui Liu

机构 * The University of Hong Kong(香港大学) Zhejiang University(浙江大学)

专题命中 长视频与时序推理 :video generation(title,abstract);long video(title,abstract);分类 cs.CV

Comments SIGGRAPH Asia 2025, Project Page: https://context-as-memory.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏