arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-07 至 2025-10-07 共收录 17 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 5 篇

2509.24008 2025-10-07 cs.CV cs.AI 79%

FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning

Haonan Ge, Yiwei Wang, Kai-Wei Chang, Hang Wu, Yujun Cai

机构 * University of California, Merced(加州大学默塞德分校) University of California, Los Angeles(加州大学洛杉矶分校) The University of Queensland(昆士兰大学)

专题命中 视频理解 :video reasoning(title);video understanding(abstract);分类 cs.CV

Comments Underreview

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04401 2025-10-07 cs.CV cs.AI 57%

Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting

Xuyang Guo, Zekai Huang, Zhenmei Shi, Zhao Song, Jiahao Zhang

机构 * Guilin University of Electronic Technology(桂林电子科技大学) The Ohio State University(俄亥俄州立大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03955 2025-10-07 cs.CV 57%

Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs

Sameep Vani, Shreyas Jena, Maitreya Patel, Chitta Baral, Somak Aditya, Yezhou Yang

机构 * Arizona State University(亚利桑那州立大学) Indian Institute of Technology, Kharagpur(印度理工学院,克拉格浦)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 17 pages, 9 figures, 6 tables. Presents TimeWarp, a synthetic preference data framework to improve temporal understanding in Video-LLMs, showing consistent gains across seven benchmarks. Includes supplementary material in the Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03706 2025-10-07 cs.RO cs.AI cs.CV cs.LG 57%

EmbodiSwap for Zero-Shot Robot Imitation Learning

Eadom Dessalene, Pavan Mantripragada, Michael Maynord, Yiannis Aloimonos

机构 * department of Computer Science, University of Maryland, College Park, MD, 20742(计算机科学系,马里兰大学, College Park, MD)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Video link: https://drive.google.com/file/d/1UccngwgPqUwPMhBja7JrXfZoTquCx_Qe/view?usp=sharing

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11336 2025-10-07 cs.CV 57%

UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks

Peiran Wu, Yunze Liu, Zhengdong Zhu, Enmin Zhou, Junxiao Shen

机构 * University of Bristol(布里斯托大学) Memories.ai Research(Memories.ai研究)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 9 篇

2412.07730 2025-10-07 cs.CV cs.AI cs.LG cs.MM 86%

STIV: Scalable Text and Image Conditioned Video Generation

Zongyu Lin, Wei Liu, Chen Chen, Jiasen Lu, Wenze Hu, Tsu-Jui Fu, Jesse Allardice, Zhengfeng Lai, Liangchen Song, Bowen Zhang, Cha Chen, Yiran Fei, Lezhi Li, Yizhou Sun, Kai-Wei Chang, Yinfei Yang

机构 * Apple(苹果公司) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 视频生成 :video generation(title,abstract);text-to-video(abstract);long video(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05093 2025-10-07 cs.CV 83%

Character Mixing for Video Generation

Tingting Liao, Chongjian Ge, Guangyi Liu, Hao Li, Yi Zhou

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫莫德·宾·扎耶德人工智能大学)

专题命中 视频生成 :video generation(title,abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03909 2025-10-07 cs.CV 83%

Generating Human Motion Videos using a Cascaded Text-to-Video Framework

Hyelin Nam, Hyojun Go, Byeongjun Park, Byung-Hoon Kim, Hyungjin Chung

机构 * EverEx University of Michigan(密歇根大学) ETH Zurich(苏黎世联邦理工学院) Yonsei University(延世大学)

专题命中 视频生成 :text-to-video(title);video generation(abstract);video diffusion(abstract);分类 cs.CV

Comments 18 pages, 7 figures, Project Page:https://hyelinnam.github.io/Cameo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11100 2025-10-07 cs.CV 83%

DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes

Jinxiu Liu, Shaoheng Lin, Yinxiao Li, Ming-Hsuan Yang

机构 * South China University of Technology(南方科技大学) Google DeepMind(谷歌DeepMind) UC Merced(加州大学默塞德分校)

专题命中 视频生成 :video generation(title,abstract);video diffusion(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04999 2025-10-07 cs.GR cs.AI cs.CV 79%

Bridging Text and Video Generation: A Survey

Nilay Kumar, Priyansh Bhandari, G. Maragatham

机构 * Department of Computational Intelligence(计算智能系) SRM Institute of Science and Technology(SRM科学与技术学院)

专题命中 视频生成 :video generation(title);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02690 2025-10-07 cs.CV cs.AI cs.LG 79%

Controllable Video Generation with Provable Disentanglement

Yifan Shen, Peiyuan Zhu, Zijian Li, Shaoan Xie, Namrata Deka, Zongfang Liu, Zeyu Tang, Guangyi Chen, Kun Zhang

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10575 2025-10-07 cs.CV 57%

A Survey of Defenses Against AI-Generated Visual Media: Detection,Disruption, and Authentication

Jingyi Deng, Chenhao Lin, Zhengyu Zhao, Shuai Liu, Zhe Peng, Qian Wang, Chao Shen

机构 * School of Cyber Science and Engineering, Xi'an Jiaotong University(西安交通大学计算机科学与工程学院) Interdisciplinary Research Center of Frontier science and technology, Xi'an Jiaotong University(西安交通大学前沿科学与技术交叉研究中心) School of Software Engineering, Xi'an Jiaotong University(西安交通大学软件工程学院) Department of Industrial and Systems Engineering, Hong Kong Polytechnic University(香港理工大学工业与系统工程系) School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted by ACM Computing Surveys

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04390 2025-10-07 cs.CV cs.AI cs.CL 57%

MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator

Xuehai He, Shijie Zhou, Thivyanth Venkateswaran, Kaizhi Zheng, Ziyu Wan, Achuta Kadambi, Xin Eric Wang

机构 * University of California, Santa Cruz(加州大学圣克ruz分校) University of California, Los Angeles(加州大学洛杉矶分校) IIT Bombay(印度理工学院班加罗尔) Microsoft(微软公司)

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14794 2025-10-07 cs.DC cs.AI cs.LG 50%

Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference

Le Chen, Dahu Feng, Erhu Feng, Yingrui Wang, Rong Zhao, Yubin Xia, Pinjie Xu, Haibo Chen

机构 * Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University(并行与分布式系统研究所,上海交通大学) Tsinghua University(清华大学) SenseTime Research(商汤科技研究院)

专题命中 视频生成 :video generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 2 篇

2507.13934 2025-10-07 cs.CV 79%

DiViD: Disentangled Video Diffusion for Static-Dynamic Factorization

Marzieh Gheisari, Auguste Genovesio

机构 * École Normale Supérieure PSL(巴黎高等师范学院)

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03886 2025-10-07 cs.AI 50%

Rare Text Semantics Were Always There in Your Diffusion Transformer

Seil Kang, Woojung Han, Dayun Ju, Seong Jae Hwang

专题命中 视频扩散模型 :text-to-video(abstract)

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频数据与评测 1 篇

2406.19568 2025-10-07 cs.CV cs.AI 70%

How Far are AI-generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach

Chirui Chang, Jiahui Liu, Zhengzhe Liu, Xiaoyang Lyu, Yi-Hua Huang, Xin Tao, Pengfei Wan, Di Zhang, Xiaojuan Qi

机构 * The University of Hong Kong(香港大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Lingnan University(岭大)

专题命中 视频数据与评测 :video generation(abstract);video diffusion(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏