arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-21 至 2025-10-21 共收录 8 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2506.22967 2025-10-21 cs.CV cs.LG cs.MM 73%

ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment

Amir Aghdam, Vincent Tao Hu, Björn Ommer

机构 * Department of Computer Science, Temple University(Temple大学计算机科学系) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 视频理解 :video understanding(abstract);video-language(abstract);分类 cs.CV、cs.MM

Comments Accepted to TMLR 2025 - Project page: https://amir-aghdam.github.io/act-align/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16209 2025-10-21 cs.CV 57%

StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales

Nyle Siddiqui, Rohit Gupta, Sirnam Swetha, Mubarak Shah

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 3 篇

2510.16833 2025-10-21 cs.CV cs.GR 79%

From Mannequin to Human: A Pose-Aware and Identity-Preserving Video Generation Framework for Lifelike Clothing Display

Xiangyu Mu, Dongliang Zhou, Jie Hou, Haijun Zhang, Weili Guan

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14588 2025-10-21 cs.CV cs.AI 79%

STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding

Zhifei Chen, Tianshuo Xu, Leyi Wu, Luozhou Wang, Dongyu Yan, Zihan You, Wenting Luo, Guo Zhang, Yingcong Chen

机构 * HKUST(GZ)(香港科技大学(珠海)) HKUST(香港科技大学) XMU(厦门大学) MIT(麻省理工学院)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

Comments Code, model, and demos can be found at https://envision-research.github.io/STANCE/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17041 2025-10-21 cs.CV cs.AI cs.LG 79%

Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM

Jaemin Kim, Bryan Sangwoo Kim, Jong Chul Ye

机构 * Graduate School of AI, KAIST(人工智能研究生院,韩国科学技术院)

专题命中 视频生成 :text-to-video(title,abstract);分类 cs.CV

Comments ICCV 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 2 篇

2510.15874 2025-10-21 cs.GR 82%

Sketch-based Fluid Video Generation Using Motion-Guided Diffusion Models in Still Landscape Images

Hao Jin, Haoran Xie

专题命中 视频扩散模型 :video generation(title,abstract);video diffusion(abstract)

Comments 2 pages, 5 figures. SIGGRAPH 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03621 2025-10-21 cs.CV 70%

DynVFX: Augmenting Real Videos with Dynamic Content

Danah Yatim, Rafail Fridman, Omer Bar-Tal, Tali Dekel

机构 * Weizmann Institute of Science(魏兹曼科学研究所)

专题命中 视频扩散模型 :video diffusion(abstract);text-to-video(abstract);分类 cs.CV

Comments Project page: https://dynvfx.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频问答 1 篇

2510.17364 2025-10-21 cs.CV cs.LG 57%

Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs

Vaggelis Dorovatas, Soroush Seifi, Gunshi Gupta, Rahaf Aljundi

机构 * Toyota Motor Europe(丰田欧洲公司) Archimedes RU, Athena RC(阿奇米德 RU、阿泰纳 RC) University of Oxford(牛津大学)

专题命中 视频问答 :long video(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏