arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-11-27 至 2025-11-27 共收录 12 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 3 篇

2511.16595 2025-11-27 cs.CV cs.AI cs.CL 88%

TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding

TimeViper: 一种用于高效长视频理解的混合Mamba-Transformer视觉语言模型

Boshen Xu, Zihan Xiao, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Jian Luan, Qin Jin

机构 * AIM3 Lab, Renmin University of China(中国人民大学人工智能实验室) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 TimeViper是一种混合Mamba-Transformer模型,通过TransV模块实现高效长视频理解,提升多模态处理能力。

Comments Project page: https://xuboshen.github.io/TimeViper; Code: https://github.com/xiaomi-research/timeviper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01533 2025-11-27 cs.CV 83%

ReasonAct: Progressive Training for Fine-Grained Video Reasoning in Small Models

ReasonAct: 为小模型在细粒度视频推理中的渐进训练

Jiaxin Liu, Zhaolu Kang

专题命中 视频理解 :video reasoning(title,abstract);video understanding(abstract);分类 cs.CV

AI总结 ReasonAct通过三阶段训练提升小模型细粒度视频推理性能,实现比基线更高的准确率和计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12336 2025-11-27 cs.CV 79%

Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding

对多模态大语言模型在视频理解中的可信度进行基准测试

Youze Wang, Zijun Chen, Ruoyu Chen, Shishen Gu, Wenbo Hu, Jiayang Liu, Yinpeng Dong, Hang Su, Jun Zhu, Meng Wang, Richang Hong

机构 * Hefei University of Technology(合肥工业大学) Tsinghua University(清华大学) Institute of Science Tokyo(东京科学研究所)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 本研究提出Trust-videoLLMs基准,评估23种视频LLMs在真实性、鲁棒性、安全性和隐私等方面的表现,揭示其在动态场景理解及现实风险缓解中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 6 篇

2511.21129 2025-11-27 cs.CV cs.GR 88%

CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion

通过统一多模态视频扩散实现可控视频生成

Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu, Jialun Liu, Hao Pan, Yuchi Huo, Rui Wang, Haibin Huang, Chi Zhang, Xuelong Li

机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院) Tsinghua University(清华大学)

专题命中 视频生成 :video generation(title,abstract);video diffusion(title);video understanding(abstract);分类 cs.CV

AI总结 CtrlVDiff通过统一多模态视频扩散模型,实现可控视频生成,支持多模态输入并提升生成的可控性和保真度。

Comments 27 pages, 18 figures, 9 tables. Project page: https://tele-ai.github.io/CtrlVDiff/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21136 2025-11-27 cs.CV cs.AI 79%

Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning

高效训练人类视频生成的熵引导优先级渐进学习

Changlin Li, Jiawei Zhang, Shuhao Liu, Sihao Lin, Zeyi Shi, Zhihui Li, Xiaojun Chang

机构 * Stanford University(斯坦福大学) North China Electric Power University(华北电力大学) University of Adelaide(阿德莱德大学) University of Technology Sydney(悉尼技术大学) University of Science and Technology of China(中国科学技术大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

AI总结 本文提出熵引导优先级渐进学习方法,通过条件熵膨胀和自适应渐进计划,高效训练扩散模型生成人类视频,实现训练速度提升和内存消耗降低。

Comments Project page: https://github.com/changlin31/Ent-Prog

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18780 2025-11-27 cs.CV cs.AI 79%

ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection

ConceptGuard:通过多模态风险检测实现文本-图像到视频生成的主动安全

Ruize Ma, Minghong Cai, Yilei Jiang, Jiaming Han, Yi Feng, Yingshui Tan, Xiaoyong Zhu, Bo Zhang, Bo Zheng, Xiangyu Yue

机构 * CUHK MMLab(香港中文大学多模态实验室) Future Lab, Alibaba Group(阿里巴巴集团未来实验室) Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

AI总结 ConceptGuard通过多模态风险检测实现文本-图像到视频生成的主动安全,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19386 2025-11-27 cs.CV cs.AI 79%

Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals

力提示:视频生成模型可以学习并泛化基于物理的控制信号

Nate Gillman, Charles Herrmann, Michael Freeman, Daksh Aggarwal, Evan Luo, Deqing Sun, Chen Sun

机构 * Brown University(布朗大学) Google DeepMind(谷歌DeepMind)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

AI总结 本文提出力提示方法,通过物理力信号生成逼真视频,利用视觉和运动先验实现物理控制信号的泛化,提升世界模型的物理真实性。

Comments Camera ready version (NeurIPS 2025). Code and interactive demos at https://force-prompting.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20462 2025-11-27 cs.CV cs.LG 70%

STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows

STARFlow-V:基于归一化流的端到端视频生成模型

Jiatao Gu, Ying Shen, Tianrong Chen, Laurent Dinh, Yuyang Wang, Miguel Angel Bautista, David Berthelot, Josh Susskind, Shuangfei Zhai

机构 * Apple(苹果公司)

专题命中 视频生成 :video generation(abstract);text-to-video(abstract);分类 cs.CV

AI总结 STARFlow-V基于归一化流提出端到端视频生成模型,具备端到端学习、鲁棒因果预测和原生似然估计等优势,实现了高质量自回归视频生成。

Comments 21 pages, 9 figures. Code and samples are available at https://github.com/apple/ml-starflow

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21475 2025-11-27 cs.CV 57%

MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices

MobileI2V: 一种适用于移动设备的快速高分辨率图像到视频生成方法

Shuai Zhang, Bao Tang, Siyuan Yu, Yueting Zhu, Jingfeng Yao, Ya Zou, Shanglin Yuan, Li Yu, Wenyu Liu, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

AI总结 MobileI2V通过轻量级扩散模型和优化策略,在移动设备上实现快速高分辨率图像到视频生成,生成速度提升10倍,质量与现有模型相当。

Comments Our Demo and code:https://github.com/hustvl/MobileI2V

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 2 篇

2511.21592 2025-11-27 cs.CV 83%

MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training

通过少量步骤的运动对抗性后训练提升视频扩散中的运动质量

Haotian Xue, Qi Chen, Zhonghao Wang, Xun Huang, Eli Shechtman, Jinrong Xie, Yongxin Chen

机构 * Adobe(Adobe公司) Georgia Tech(佐治亚理工学院)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

AI总结 MoGAN通过运动对抗性后训练提升视频扩散模型的运动质量,无需奖励模型或人类偏好数据,在多个基准测试中显著提高了运动真实感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18255 2025-11-27 cs.CV 57%

Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization

基于扩散噪声优化的连续流序列适应视频预测

Sina Mokhtarzadeh Azar, Emad Bahrami, Enrico Pallotta, Gianpiero Francesca, Radu Timofte, Juergen Gall

机构 * University of Bonn(波恩大学) Toyota Motor Europe(丰田欧洲公司) University of Wuerzburg(乌尔姆大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究所)

专题命中 视频扩散模型 :long video(abstract);分类 cs.CV

AI总结 本文提出了一种基于扩散噪声优化的连续流序列适应视频预测方法,通过在推理过程中细化扩散噪声以提升视频预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频问答 1 篇

2511.17490 2025-11-27 cs.CV 79%

Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination

Video-R4:通过视觉沉思强化文本丰富的视频推理

Yolo Y. Tang, Daiki Shimada, Hang Hua, Chao Huang, Jing Bi, Rogerio Feris, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) Sony Group Corporation(索尼集团) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)

专题命中 视频问答 :video reasoning(title,abstract);分类 cs.CV

AI总结 Video-R4通过视觉沉思机制提升文本丰富视频的推理能力,采用多阶段学习框架实现像素基础的多模态推理。

详情

展开后加载摘要…

URL PDF HTML 收藏