arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-11-17 至 2025-11-17 共收录 9 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1 篇

2511.04668 2025-11-17 cs.CV 79%

SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding

Ellis Brown, Arijit Ray, Ranjay Krishna, Ross Girshick, Rob Fergus, Saining Xie

机构 * New York University(纽约大学) Boston University(波士顿大学) AllenAI Vercept

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments Project page: https://ellisbrown.github.io/sims-v

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 3 篇

2411.16657 2025-11-17 cs.CV cs.AI cs.CL 83%

DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation

Zun Wang, Jialu Li, Han Lin, Jaehong Yoon, Mohit Bansal

专题命中 视频生成 :video generation(title,abstract);text-to-video(abstract);分类 cs.CV

Comments AAAI 2026, Project website: https://zunwang1.github.io/DreamRunner

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10202 2025-11-17 cs.CL cs.CV 79%

Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval

Shubhashis Roy Dipta, Francis Ferraro

机构 * Department of Computer Science and Electrical Engineering University of Maryland Baltimore County(计算机科学与电气工程系马里兰大学巴尔的摩县)

专题命中 视频生成 :text-to-video(title,abstract);分类 cs.CV

Comments Accepted in IJCNLP-AACL 2025 (also presented in MAGMAR 2025 at ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25187 2025-11-17 cs.CV 74%

FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation

Yunyang Ge, Xinhua Cheng, Chengshu Zhao, Xianyi He, Shenghai Yuan, Bin Lin, Bin Zhu, Li Yuan

机构 * Peking University, Shenzhen Graduate School(北京大学深圳研究生院) Peng Cheng Laboratory(鹏城实验室) Rabbitpre AI

专题命中 视频生成 :video generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 2 篇

2511.11062 2025-11-17 cs.CV cs.AI 70%

LiteAttention: A Temporal Sparse Attention for Diffusion Transformers

Dor Shmilovich, Tony Wu, Aviad Dahan, Yuval Domb

专题命中 视频扩散模型 :video generation(abstract);video diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11213 2025-11-17 cs.CV 57%

RealisticDreamer: Guidance Score Distillation for Few-shot Gaussian Splatting

Ruocheng Wu, Haolan He, Yufei Wang, Zhihao Li, Bihan Wen

机构 * University of Electronic Science and Technology of China(电子科技大学) Nanyang Technological University(南洋理工大学)

专题命中 视频扩散模型 :video diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频问答 1 篇

2509.18041 2025-11-17 cs.CV 79%

NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning

Sahil Shah, S P Sharan, Harsh Goel, Minkyu Choi, Mustafa Munir, Manvik Pasula, Radu Marculescu, Sandeep Chinchali

专题命中 视频问答 :video understanding(title);long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 长视频与时序推理 1 篇

2511.10866 2025-11-17 cs.CV cs.AI 57%

Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling

Seoik Jung, Taekyung Song, Yangro Lee, Sungjun Lee

机构 * PIA-SPACE Inc.(PIA-SPACE公司)

专题命中 长视频与时序推理 :long video(abstract);分类 cs.CV

Comments 5 pages, 2 figures. Accepted paper for the IEIE (Institute of Electronics and Information Engineers) Fall Conference 2025. Presentation on Nov 27, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 视频数据与评测 1 篇

2511.11002 2025-11-17 cs.CV 83%

EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation

Zongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He, Zeyu Wang

专题命中 视频数据与评测 :video understanding(title);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments 15 pages, 12 figures. Accepted as an Oral presentation at AAAI 2026. For code and dataset, see https://zane-zyqiu.github.io/EmoVid

详情

展开后加载摘要…

URL PDF HTML 收藏