arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-08-08 至 2025-08-08 共收录 6 信号源:cs.CV, eess.IV, cs.MM

1. 视频扩散模型 3 篇

2502.15894 2025-08-08 cs.CV 85%

RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers

Min Zhao, Guande He, Yixiao Chen, Hongzhou Zhu, Chongxuan Li, Jun Zhu

机构 * Dept. of Comp. Sci. \& Tech., BNRist Center, THU-Bosch ML Center, Tsinghua University. The University of Texas at Austin. Gaoling School of Artificial Intelligence Renmin University of China Beijing, China. Beijing Key Laboratory of Research on Large Models Engineering Research Center of Next-Generation Intelligent Search Pazhou Laboratory (Huangpu)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);long video(abstract);分类 cs.CV

Comments ICML 2025. Project page: https://riflex-video.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23287 2025-08-08 cs.CV 70%

ReferEverything: Towards Segmenting Everything We Can Speak of in Videos

Anurag Bagchi, Zhipeng Bao, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Toyota Research Institute(丰田研究 institute)

专题命中 视频扩散模型 :video generation(abstract);video diffusion(abstract);分类 cs.CV

Comments Project page at https://refereverything.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01603 2025-08-08 cs.CV 57%

DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation

Yue-Jiang Dong, Wang Zhao, Jiale Xu, Ying Shan, Song-Hai Zhang

机构 * Tsinghua University(清华大学) ARC Lab, Tencent PCG(腾讯PCG实验室)

专题命中 视频扩散模型 :long video(abstract);分类 cs.CV

Comments Accepted by ICCV 2025; Project Homepage: https://yuejiangdong.github.io/depthsync

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 长视频与时序推理 1 篇

2505.08444 2025-08-08 cs.RO 50%

Extracting Visual Plans from Unlabeled Videos via Symbolic Guidance

Wenyan Yang, Ahmet Tikna, Yi Zhao, Yuying Zhang, Luigi Palopoli, Marco Roveri, Joni Pajarinen

机构 * Department of Electrical Engineering and Automation, Aalto University(艾尔沃斯大学电气工程与自动化系) Department of Engineering and Computer Science, University of Trento(特伦托大学工程与计算机科学系)

专题命中 长视频与时序推理 :video generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频数据与评测 2 篇

2508.05527 2025-08-08 cs.CV 57%

AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety

Adi Levi, Or Levi, Sardhendu Mishra, Jonathan Morra

机构 * Zefr Inc(Zefr公司)

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments Accepted to the Computer Vision in Advertising and Marketing (CVAM) workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09994 2025-08-08 cs.CV 57%

TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs

Yunxiao Wang, Meng Liu, Wenqi Liu, Xuemeng Song, Bin Wen, Fan Yang, Tingting Gao, Di Zhang, Guorui Zhou, Liqiang Nie

机构 * Shandong University(山东大学) Shandong Jianzhu University(山东建筑大学) City University of Hong Kong(香港城市大学) Kuaishou Technology(快手科技)

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏