arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-09-24 至 2025-09-24 共收录 9 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2409.02889 2025-09-24 cs.CL cs.AI cs.CV cs.MM 62%

LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture

Xidong Wang, Dingjie Song, Shunian Chen, Junyin Chen, Zhenyang Cai, Chen Zhang, Lichao Sun, Benyou Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Lehigh University(莱斯利大学) Meituan(美团)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19245 2025-09-24 cs.CV 57%

ConViS-Bench: Estimating Video Similarity Through Semantic Concepts

Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero, Yiming Wang, Elisa Ricci, Paolo Rota

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 2 篇

2506.00329 2025-09-24 cs.LG cs.AI cs.CV 88%

Foresight: Adaptive Layer Reuse for Accelerated and High-Quality Text-to-Video Generation

Muhammad Adnan, Nithesh Kurella, Akhil Arunkumar, Prashant J. Nair

机构 * The University of British Columbia(不列颠哥伦比亚大学) d-Matrix

专题命中 视频生成 :video generation(title,abstract);text-to-video(title,abstract);分类 cs.CV

Comments Accepted at the 39th Conference on Neural Information Processing Systems (NeurIPS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19222 2025-09-24 cs.LG 78%

Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models

Julien Delavande, Regis Pierrard, Sasha Luccioni

机构 * Hugging Face

专题命中 视频生成 :text-to-video(title,abstract)

Comments 10 pages. Accepted as an oral presentation at the NeurIPS 2025 NextVid Workshop (San Diego, December 6, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 3 篇

2508.03485 2025-09-24 cs.CV 83%

LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation

Lianwei Yang, Haokun Lin, Tianchen Zhao, Yichen Wu, Hongyu Zhu, Ruiqi Xie, Zhenan Sun, Yu Wang, Qingyi Gu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院) Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)

专题命中 视频扩散模型 :video generation(title,abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19296 2025-09-24 cs.CV cs.GR 79%

Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation

Sherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang, Yifeng Jiang, Haithem Turki, Andrea Tagliasacchi, David B. Lindell, Zan Gojcic, Sanja Fidler, Huan Ling, Jun Gao, Xuanchi Ren

机构 * NVIDIA University of Toronto(多伦多大学) Vector Institute(向量研究所) Simon Fraser University(西蒙弗雷泽大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

Comments Project Page: https://research.nvidia.com/labs/toronto-ai/lyra/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06028 2025-09-24 cs.CV 57%

SparseDiT: Token Sparsification for Efficient Diffusion Transformer

Shuning Chang, Pichao Wang, Jiasheng Tang, Fan Wang, Yi Yang

机构 * Zhejiang University(浙江大学) Damo Academy, Alibaba Group(阿里达摩院) Hupan Lab(虎斑实验室)

专题命中 视频扩散模型 :video generation(abstract);分类 cs.CV

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 长视频与时序推理 1 篇

2503.04130 2025-09-24 cs.CV 89%

STORM: Token-Efficient Long Video Understanding for Multimodal LLMs

Jindong Jiang, Xiuyu Li, Zhijian Liu, Muyang Li, Guo Chen, Zhiqi Li, De-An Huang, Guilin Liu, Zhiding Yu, Kurt Keutzer, Sungjin Ahn, Jan Kautz, Hongxu Yin, Yao Lu, Song Han, Wonmin Byeon

机构 * NVIDIA(英伟达) Rutgers University(罗格斯大学) UC Berkeley(加州大学伯克利分校) MIT(麻省理工学院) Nanjing University(南京大学) KAIST(韩国科学技术院)

专题命中 长视频与时序推理 :video understanding(title,abstract);long video(title,abstract);video reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 视频数据与评测 1 篇

2505.15173 2025-09-24 cs.CV cs.AI 57%

AvatarShield: Visual Reinforcement Learning for Human-Centric Synthetic Video Detection

Zhipei Xu, Xuanyu Zhang, Qing Huang, Xing Zhou, Jian Zhang

专题命中 视频数据与评测 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏