arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-09-23 至 2025-09-23 共收录 14 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2509.16810 2025-09-23 cs.AI 78%

Automated Procedural Analysis via Video-Language Models for AI-assisted Nursing Skills Assessment

Shen Chang, Dennis Liu, Renran Tian, Kristen L. Swartzell, Stacie L. Klingler, Amy M. Nagle, Nan Kong

机构 * Weldon School of Biomedical Engineering, Purdue University(普渡大学生物医学工程学院) Department of Industrial and Operations Engineering, University of Michigan(密歇根大学工业与运作工程系) Edward P. Fitts Department of Industrial and Systems Engineering, North Carolina State University(北卡罗来纳州立大学工业与系统工程系) School of Nursing, Purdue University(普渡大学护理学院)

专题命中 视频理解 :video-language(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16560 2025-09-23 cs.CV 57%

Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization

Ji Soo Lee, Byungoh Ko, Jaewon Cho, Howoong Lee, Jaewoon Byun, Hyunwoo J. Kim

机构 * Korea University(韩国大学) Hanwha Vision(翰威英航) KAIST(韩国科学技术院)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 6 篇

2509.16956 2025-09-23 cs.CV 88%

VidCLearn: A Continual Learning Approach for Text-to-Video Generation

Luca Zanchetta, Lorenzo Papa, Luca Maiano, Irene Amerini

机构 * Sapienza University of Rome, Italy(罗马大学萨皮恩扎) ESA philab(欧洲航天局实验室)

专题命中 视频生成 :video generation(title,abstract);text-to-video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12667 2025-09-23 cs.CV 86%

Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking

Zihan Su, Xuerui Qiu, Hongbin Xu, Tangyu Jiang, Junhao Zhuang, Chun Yuan, Ming Li, Shengfeng He, Fei Richard Yu

机构 * Tsinghua University(清华大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) South China University of Technology(华南理工大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室) Singapore Management University(新加坡管理学院)

专题命中 视频生成 :video generation(title,abstract);text-to-video(title);分类 cs.CV

Comments Safa-Sora is accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14899 2025-09-23 cs.CV 79%

Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation

Chenjie Cao, Jingkai Zhou, Shikai Li, Jingyun Liang, Chaohui Yu, Fan Wang, Xiangyang Xue, Yanwei Fu

机构 * Alibaba DAMO Academy(阿里巴巴达摩院) Fudan University(复旦大学) Hupan Lab(虎扑实验室) Shanghai Innovation Institute(上海创新研究院)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

Comments Project page: https://github.com/ewrfcas/Uni3C. Accepted by Siggraph Asian 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17190 2025-09-23 cs.CV cs.AI 74%

Echo-Path: Pathology-Conditioned Echo Video Generation

Kabir Hamzah Muhammad, Marawan Elbatel, Yi Qin, Xiaomeng Li

机构 * Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments 10 pages, 3 figures, MICCAI-AMAI2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15564 2025-09-23 cs.CV 61%

Show-o2: Improved Native Unified Multimodal Models

Jinheng Xie, Zhenheng Yang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室) ByteDance(字节跳动)

专题命中 视频生成 :video generation(abstract);分类 cs.CV;video understanding(comments)

Comments NeurIPS 2025. (v3: update to include video understanding, OneIG, and more ablation study results)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21385 2025-09-23 cs.HC 50%

Dynamic Vision from EEG Brain Recordings, How much does EEG know?

Prajwal Singh, Anupam Sharma, Pankaj Pandey, Krishna Miyapuram, Shanmuganathan Raman

专题命中 视频生成 :video generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 1 篇

2509.17985 2025-09-23 cs.GR 86%

VideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion Models

Geonung Kim, Janghyeok Han, Sunghyun Cho

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(title)

Comments Project page: https://kimgeonung.github.io/VideoFrom3D/

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 动作与事件理解 2 篇

2502.07239 2025-09-23 cs.CV cs.AI 79%

Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation

Pinxin Liu, Pengfei Zhang, Hyeongwoo Kim, Pablo Garrido, Ari Shapiro, Kyle Olszewski

机构 * University of Rochester(罗切斯特大学) University of California Irvine(加州大学尔湾分校) Imperial College(帝国理工学院)

专题命中 动作与事件理解 :video generation(title,abstract);分类 cs.CV

Comments Accepted to ACM MM 2025. Project Page: https://andypinxinliu.github.io/Contextual-Gesture/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17864 2025-09-23 cs.CV 57%

ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular Videos

Shi Chen, Erik Sandström, Sandro Lombardi, Siyuan Li, Martin R. Oswald

机构 * ETH Zürich(苏黎世联邦理工学院) Google(谷歌) University of Amsterdam(阿姆斯特丹大学)

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 长视频与时序推理 1 篇

2509.17500 2025-09-23 cs.CV 57%

SAMSON: 3rd Place Solution of LSVOS 2025 VOS Challenge

Yujie Xie, Hongyang Zhang, Zhihui Liu, Shihai Ruan

机构 * Truesight Research(Truesight研究机构) School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(科学与工程学院,香港中文大学(深圳))

专题命中 长视频与时序推理 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 视频数据与评测 2 篇

2509.16438 2025-09-23 cs.CV cs.CL 57%

AutoArabic: A Three-Stage Framework for Localizing Video-Text Retrieval Benchmarks

Mohamed Eltahir, Osamah Sarraj, Abdulrahman Alfrihidi, Taha Alshatiri, Mohammed Khurd, Mohammed Bremoo, Tanveer Hussain

专题命中 视频数据与评测 :text-to-video(abstract);分类 cs.CV

Comments Accepted at ArabicNLP 2025 (EMNLP 2025 workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04415 2025-09-23 cs.CL 50%

MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind

Emilio Villa-Cueva, S M Masrur Ahmed, Rendi Chevi, Jan Christian Blaise Cruz, Kareem Elzeky, Fermin Cristobal, Alham Fikri Aji, Skyler Wang, Rada Mihalcea, Thamar Solorio

机构 * MBZUAI University of Houston(德克萨斯大学休斯顿分校) McGill University(麦吉尔大学) University of Michigan(密歇根大学)

专题命中 视频数据与评测 :long video(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏