arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-07-29 至 2025-07-29 共收录 17 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2507.20518 2025-07-29 cs.CV cs.MM 62%

T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval

Yili Li, Gang Xiong, Gaopeng Gou, Xiangyan Qu, Jiamin Zhuang, Zhen Li, Junzheng Shi

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 视频理解 :text-to-video(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20939 2025-07-29 cs.CV 57%

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li, Teng Wang, Junfu Pu, Yizhuo Li, Lu Qiu, Jin Ma, Lisheng Duan, Xinyu Zuo, Jinwen Luo, Weibo Gu, Zexuan Li, Xiaojing Zhang, Yangyu Tao, Han Hu, Di Wang, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Search Application Department, Tencent CSIG(腾讯CSIG搜索应用部门) Tencent Hunyuan(腾讯文生视频) Big Data Platform Department, Tencent PCG(腾讯PCG大数据平台部门)

专题命中 视频理解 :video reasoning(abstract);分类 cs.CV

Comments Project Page: https://tencentarc.github.io/posts/arc-video-announcement/

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 6 篇

2504.16907 2025-07-29 cs.CV cs.AI 86%

BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation

Ruotong Wang, Mingli Zhu, Jiarong Ou, Rui Chen, Xin Tao, Pengfei Wan, Baoyuan Wu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 视频生成 :text-to-video(title,abstract);video generation(title);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19836 2025-07-29 cs.GR cs.AI cs.CV cs.MM cs.SD 81%

ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion

Xuanchen Wang, Heng Wang, Weidong Cai

机构 * School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV、cs.MM

Comments 10 pages, 5 figures, accepted by the 33rd ACM International Conference on Multimedia (ACM MM 2025), demo page: https://choreomuse.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20855 2025-07-29 cs.CV 57%

Compositional Video Synthesis by Temporal Object-Centric Learning

Adil Kaan Akan, Yucel Yemez

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 12+21 pages, submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18945 2025-07-29 cs.CV cs.AI cs.LG cs.RO 57%

Aether: Geometric-Aware Unified World Modeling

Aether Team, Haoyi Zhu, Yifan Wang, Jianjun Zhou, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Chunhua Shen, Jiangmiao Pang, Tong He

机构 * USTC Shanghai AI Lab(USTC上海人工智能实验室) SII SJTU(SJTU信息研究所) ZJU FDU(浙江大学福州大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Project Page: https://aether-world.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06438 2025-07-29 cs.CV 57%

Qffusion: Controllable Portrait Video Editing via Quadrant-Grid Attention Learning

Maomao Li, Lijian Lin, Yunfei Liu, Ye Zhu, Yu Li

机构 * International Digital Economy Academy (IDEA)(国际数字经济学院)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 19 pages

Journal ref TVCG 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12789 2025-07-29 cs.CV 57%

Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting

Haoyu Zhao, Hao Wang, Xingyue Zhao, Hao Fei, Hongqiu Wang, Chengjiang Long, Hua Zou

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Wuhan National Laboratory for Optoelectronics, Huazhong University of Science and Technology(华中科技大学光电研究院) Meta Reality Lab(Meta现实实验室) Xi’an Jiao Tong University(西安交通大学) National University of Singapore(新加坡国立大学) The Department of Systems Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学系统枢纽部门(广州))

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 5 篇

2504.03140 2025-07-29 cs.CV 79%

Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models

Xuran Ma, Yexin Liu, Yaofu Liu, Xianfeng Wu, Mingzhe Zheng, Zihao Wang, Ser-Nam Lim, Harry Yang

机构 * Hong Kong University of Science and Technology(香港科技大学) Everlyn AI University of Central Florida(佛罗里达中央大学)

专题命中 视频扩散模型 :video diffusion(title);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19789 2025-07-29 cs.CV 79%

TransFlow: Motion Knowledge Transfer from Video Diffusion Models to Video Salient Object Detection

Suhwan Cho, Minhyeok Lee, Jungho Lee, Sunghun Yang, Sangyoun Lee

机构 * GenGenAI Yonsei University(延世大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

Comments ICCVW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20158 2025-07-29 cs.CV 57%

AnimeColor: Reference-based Animation Colorization with Diffusion Transformers

Yuhong Zhang, Liyao Wang, Han Wang, Danni Wu, Zuzeng Lin, Feng Wang, Li Song

机构 * Shanghai Jiao Tong University(上海交通大学) Tianjin University(天津大学) Communication University of China(中国通信大学)

专题命中 视频扩散模型 :video diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08714 2025-07-29 cs.CV cs.AI 57%

Versatile Multimodal Controls for Expressive Talking Human Animation

Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li, Zixin Zhu, Sanping Zhou, Ming Yang, Le Wang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University, Ant Group(人机混合增强智能国家级实验室,西安交通大学,蚂蚁集团) National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University(人机混合增强智能国家级实验室,西安交通大学) University at Buffalo(布法罗大学)

专题命中 视频扩散模型 :video diffusion(abstract);分类 cs.CV

Comments Accepted by ACM MM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03456 2025-07-29 cs.CV 57%

LM-Gaussian: Boost Sparse-view 3D Gaussian Splatting with Large Model Priors

Hanyang Yu, Xiaoxiao Long, Ping Tan

专题命中 视频扩散模型 :video diffusion(abstract);分类 cs.CV

Comments Project page: https://hanyangyu1021.github.io/lm-gaussian.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频问答 1 篇

2507.19599 2025-07-29 cs.CV 70%

Object-centric Video Question Answering with Visual Grounding and Referring

Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Weidi Xie, Stratis Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) SAI, Shanghai Jiao Tong University(上海交通大学SAI研究所) Xiaohongshu Inc(小红书公司)

专题命中 视频问答 :video understanding(abstract);video reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 动作与事件理解 1 篇

2504.13092 2025-07-29 cs.CV 57%

EventVAD: Training-Free Event-Aware Video Anomaly Detection

Yihua Shao, Haojin He, Sijie Li, Siyu Chen, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma, Xiaochen Wang, Hao Tang, Yan Wang, Shuyan Li

机构 * Peking University(北京大学) Guangdong University of Technology(广东工业大学) The University of Sheffield(谢菲尔德大学) University of Science and Technology Beijing(北京科技大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Nanjing University(南京大学) University of Trento(特伦特大学) Queen's University Belfast(贝尔法斯特女王大学)

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

Comments Paper was accepted by ACM MM 2025; Code: https://github.com/YihuaJerry/EventVAD

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 长视频与时序推理 1 篇

2507.19754 2025-07-29 cs.CV 57%

Latest Object Memory Management for Temporally Consistent Video Instance Segmentation

Seunghun Lee, Jiwan Seo, Minwoo Choi, Kiljoon Han, Jaehoon Jeong, Zane Durante, Ehsan Adeli, Sang Hyun Park, Sunghoon Im

机构 * DGIST, Daegu, Republic of Korea(韩国大邱科学技术院) Stanford University(斯坦福大学)

专题命中 长视频与时序推理 :long video(abstract);分类 cs.CV

Comments ICCV 2025. Code: https://github.com/Seung-Hun-Lee/LOMM

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 视频数据与评测 1 篇

2507.20368 2025-07-29 cs.CV cs.MM 62%

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation

Shuolin Xu, Bingyuan Wang, Zeyu Cai, Fangteng Fu, Yue Ma, Tongyi Lee, Hongchuan Yu, Zeyu Wang

机构 * National Centre for Computer Animation, Bournemouth University(伯恩茅斯大学计算机动画国家中心) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Hong Kong University of Science and Technology(香港科技大学) Department of Computer Science and Information Engineering, National Cheng Kung University(国立成功大学计算机科学与信息工程系)

专题命中 视频数据与评测 :video generation(abstract);分类 cs.CV、cs.MM

Comments 8 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏