arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-02 至 2025-10-02 共收录 5 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 5 篇

2510.00806 2025-10-02 cs.CV 79%

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心、自动化研究所、中国科学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室、深圳中国) School of Artificial Intelligence, University of Chinese Academy of Science, Beijing, China(人工智能学院、中国科学院大学、北京中国) Wuhan AI Research, Wuhan, China(武汉人工智能研究、武汉中国) MAPLE Lab, Westlake University(MAPLE实验室、西湖大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00050 2025-10-02 cs.MM cs.AI cs.CV cs.SD eess.AS 62%

Object-AVEdit: An Object-level Audio-Visual Editing Model

Youquan Fu, Ruiyang Si, Hongfa Wang, Dongzhan Zhou, Jiacheng Sun, Ping Luo, Di Hu, Hongyuan Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI)(人工智能研究院) China Telecom(中国电信) Gaoling School of Artificial Intelligence(东城区人工智能学院) Renmin University of China(中国人民大学) Beijing University of Posts and Telecommunications(北京邮电大学) Tencent Data Platform(腾讯数据平台) Tsinghua University(清华大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Huawei Noah’s Ark Lab(华为诺亚实验室) The University of Hong Kong(香港大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10958 2025-10-02 cs.LG cs.AI cs.CV cs.NE cs.PF 57%

SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization

Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, Jianfei Chen

机构 * Dept. of Comp. Sci. and Tech., Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint ML Center, Tsinghua University(计算机科学与技术系,人工智能研究院,BNRist中心,THBI实验室,清华-博世联合机器学习中心,清华大学) Institute for Interdisciplinary Information Sciences, Tsinghua University(交叉信息学院,清华大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments @inproceedings{zhang2024sageattention2, title={Sageattention2: Efficient attention with thorough outlier smoothing and per-thread int4 quantization}, author={Zhang, Jintao and Huang, Haofeng and Zhang, Pengle and Wei, Jia and Zhu, Jun and Chen, Jianfei}, booktitle={International Conference on Machine Learning (ICML)}, year={2025} }

Journal ref Proceedings of the 42nd International Conference on Machine Learning, PMLR 267, 2025 (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14685 2025-10-02 cs.CV 57%

DACoN: DINO for Anime Paint Bucket Colorization with Any Number of Reference Images

Kazuma Nagata, Naoshi Kaneko

机构 * Tokyo Denki University(东京电讯大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. v2: Added results on the subset used by the baseline for consistency; full test set results are also reported (Tables 1 and 2)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02367 2025-10-02 cs.LG 50%

SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration

Jintao Zhang, Jia Wei, Haofeng Huang, Pengle Zhang, Jun Zhu, Jianfei Chen

机构 * Dept. of Comp. Sci. & Tech.(计算机科学与技术系) Institute for AI(人工智能研究院) BNRist Center(BNRist中心) Tsinghua-Bosch Joint ML Center(清华大学-博世联合机器学习中心) THBI Lab(THBI实验室) Tsinghua University(清华大学)

专题命中 视频生成 :video generation(abstract)

Comments @inproceedings{zhang2025sageattention, title={SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration}, author={Zhang, Jintao and Wei, Jia and Zhang, Pengle and Zhu, Jun and Chen, Jianfei}, booktitle={International Conference on Learning Representations (ICLR)}, year={2025} }

Journal ref The Thirteenth International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏