arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-02-02 至 2026-02-02 共收录 5 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 5 篇

2601.10214 2026-02-02 cs.CV cs.GR 83%

Beyond Inpainting: Unleash 3D Understanding for Precise Camera-Controlled Video Generation

超越修复:释放3D理解以实现精确的相机控制视频生成

Dong-Yu Chen, Yixin Guo, Shuojin Yang, Tai-Jiang Mu, Shi-Min Hu

机构 * BNRist, Department of Computer Science and Technology, Tsinghua University(BNRist,计算机科学与技术系,清华大学)

专题命中 视频生成 :video generation(title,abstract);video diffusion(abstract);分类 cs.CV

AI总结 DepthDirector通过双流条件机制和LoRA适配器实现精确摄像机控制,提升视频生成的3D理解和视觉质量。

Comments Project page: https://eleanor6725.github.io/DepthDirector/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05179 2026-02-02 cs.CV 83%

FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation

FlashVideo: 为高效高分辨率视频生成实现流畅的细节保真度

Shilong Zhang, Wenbo Li, Shoufa Chen, Chongjian GE, Peize Sun, Yifu Zhang, Yi Jiang, Zehuan Yuan, Bingyue Peng, Ping Luo

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学) ByteDance(字节跳动)

专题命中 视频生成 :video generation(title,abstract);text-to-video(abstract);分类 cs.CV

AI总结 FlashVideo通过两阶段框架在高效高分辨率视频生成中实现流畅细节保真度,优化计算效率并提升商业应用潜力。

Comments Model and Weight: https://github.com/FoundationVision/FlashVideo

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19488 2026-02-02 cs.CV 79%

Entropy-Guided k-Guard Sampling for Long-Horizon Autoregressive Video Generation

熵引导的k-guard采样用于长 Horizon 自回归视频生成

Yizhao Han, Tianxing Shi, Zhao Wang, Zifan Xu, Zhiyuan Pu, Mingxiao Li, Qian Zhang, Wei Yin, Xiao-Xiao Long

机构 * Nanjing University(南京大学) Horizon Robotics China Mobile(中国移动)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

AI总结 熵引导的k-guard采样用于长Horizon自回归视频生成,通过自适应调整令牌候选集大小以提升视频生成的质量和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14256 2026-02-02 cs.CV 79%

Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning

身份保持的多人类视频生成优化:通过强化学习进行优化

Xiangyu Meng, Zixian Zhang, Zhenghao Zhang, Junchao Liao, Long Qin, Weizhi Wang

机构 * Alibaba Group(阿里巴巴集团) Fudan University(复旦大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

AI总结 Identity-GRPO通过强化学习优化多人类身份保持的视频生成,显著提升一致性指标。

Comments Our project and code are available at https://ali-videoai.github.io/identity_page, https://github.com/alibaba/identity-grpo

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02725 2026-02-02 cs.AI 50%

Advances in Artificial Intelligence: A Review for the Creative Industries

人工智能进展:面向创意产业的综述

Nantheera Anantrasirichai, Fan Zhang, David Bull

机构 * Visual Information Laboratory, University of Bristol, Bristol, UK(布里斯托大学视觉信息实验室)

专题命中 视频生成 :video generation(abstract)

AI总结 本文综述了自2022年以来人工智能在创意产业中的进展,探讨了生成式AI、大语言模型和扩散模型等技术对创意生产流程的影响,并分析了人类与AI协作的新趋势及面临的挑战。

Comments This is an updated review of our previous paper (see https://doi.org/10.1007/s10462-021-10039-7), and has been accepted by Artificial Intelligence Review journal

详情

展开后加载摘要…

URL PDF HTML 收藏