arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-09-03 至 2025-09-03 共收录 18 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 5 篇

2508.21496 2025-09-03 cs.CV cs.AI 88%

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

Hao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang, Xuanyu Zheng, Yepeng Tang, Dahua Lin, Lewei Lu

机构 * Sensetime(秒氏科技)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09146 2025-09-03 cs.CV cs.MM 85%

Generative Frame Sampler for Long Video Understanding

Linli Yao, Haoning Wu, Kun Ouyang, Yuanxing Zhang, Caiming Xiong, Bei Chen, Xu Sun, Junnan Li

机构 * National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机学院,北京大学) Peking University(北京大学) Salesforce Research(Salesforce 研究) Independent Researcher(独立研究者)

专题命中 视频理解 :video understanding(title);long video(title);分类 cs.CV、cs.MM

Comments ACL 2025 Findings. Code: https://github.com/yaolinli/GenS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13919 2025-09-03 cs.CV cs.AI cs.CL cs.LG cs.RO 79%

Temporal Preference Optimization for Long-Form Video Understanding

Rui Li, Xiaohan Wang, Yuhui Zhang, Orr Zohar, Zeyu Wang, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00357 2025-09-03 cs.CV cs.AI cs.LG 79%

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Zhen Chen, Xingjian Luo, Kun Yuan, Jinlin Wu, Danny T. M. Chan, Nassir Navab, Hongbin Liu, Zhen Lei, Jiebo Luo

机构 * Hong Kong Institute of Science & Innovation(香港科学与工业创新研究院) CAMP, Technische Universität München(CAMP,慕尼黑技术大学) Department of Surgery, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院外科部)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00723 2025-09-03 cs.AI cs.MM 57%

OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination

Junzhe Chen, Tianshu Zhang, Shiyu Huang, Yuwei Niu, Chao Sun, Rongzhou Zhang, Guanyu Zhou, Lijie Wen, Xuming Hu

机构 * Tsinghua University(清华大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) OpenRL Chongqing University(重庆大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 5 篇

2509.01362 2025-09-03 cs.CV cs.MM 90%

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement

Jiayi Gao, Changcheng Hua, Qingchao Chen, Yuxin Peng, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院)

专题命中 视频生成 :video generation(title,abstract);text-to-video(title,abstract);video diffusion(abstract);分类 cs.CV、cs.MM

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20491 2025-09-03 cs.CV cs.CL cs.LG 88%

VPO: Aligning Text-to-Video Generation Models with Prompt Optimization

Jiale Cheng, Ruiliang Lyu, Xiaotao Gu, Xiao Liu, Jiazheng Xu, Yida Lu, Jiayan Teng, Zhuoyi Yang, Yuxiao Dong, Jie Tang, Hongning Wang, Minlie Huang

机构 * The Conversational Artificial Intelligence (CoAI) Group, Tsinghua University(清华大学对话人工智能(CoAI)小组) Zhipu AI(智谱AI) The Knowledge Engineering Group (KEG), Tsinghua University(清华大学知识工程小组(KEG))

专题命中 视频生成 :video generation(title,abstract);text-to-video(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01277 2025-09-03 cs.AI 82%

Communicative Agents for Slideshow Storytelling Video Generation based on LLMs

Jingxing Fan, Jinrong Shen, Yusheng Yao, Shuangqing Wang, Qian Wang, Yuling Wang

专题命中 视频生成 :video generation(title,abstract);text-to-video(abstract)

Comments 8 pages, 8 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01232 2025-09-03 cs.CV 57%

FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework

Lingzhou Mu, Qiang Wang, Fan Jiang, Mengchao Wang, Yaqi Fan, Mu Xu, Kai Zhang

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments https://fantasy-amap.github.io/fantasy-hsi/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00403 2025-09-03 cs.CV 57%

DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective

Yushuo Chen, Ruizhi Shao, Youxin Pang, Hongwen Zhang, Xinyi Wu, Rihui Wu, Yebin Liu

机构 * Tsinghua University(清华大学) Beijing Normal University(北京师范大学) Honor Device Co., Ltd(荣誉设备有限公司)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 1 篇

2509.00843 2025-09-03 cs.CV cs.AI 79%

Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion

Xueyang Kang, Zhengkang Xiang, Zezheng Zhang, Kourosh Khoshelham

机构 * University of Melbourne(墨尔本大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

Comments 26 pages, 30 figures, 2025 ACM Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 动作与事件理解 2 篇

2503.10684 2025-09-03 cs.CV cs.AI 57%

Open-World Skill Discovery from Unsegmented Demonstrations

Jingwen Deng, Zihao Wang, Shaofei Cai, Anji Liu, Yitao Liang

机构 * Peking University(北京大学) University of California, Los Angeles(美国加州大学洛杉矶分校)

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01954 2025-09-03 cs.SI cs.CL cs.CY cs.ET cs.LG 50%

Content and Engagement Trends in COVID-19 YouTube Videos: Evidence from the Late Pandemic

Nirmalya Thakur, Madeline D Hartel, Lane Michael Boden, Dallas Enriquez, Boston Joyner Ricks

专题命中 动作与事件理解 :long video(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 长视频与时序推理 2 篇

2507.23575 2025-09-03 cs.CV 57%

Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation

Sobhan Asasi, Mohamed Ilyas Lakhal, Ozge Mercanoglu Sincan, Richard Bowden

机构 * Center for Vision, Speech and Signal Processing (CVSSP) University of Surrey(视觉、语音和信号处理中心(CVSSP)大学学院)

专题命中 长视频与时序推理 :long video(abstract);分类 cs.CV

Comments Accepted at BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13138 2025-09-03 cs.CV 57%

STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation

Jiamin Wang, Yichen Yao, Xiang Feng, Hang Wu, Yaming Wang, Qingqiu Huang, Yuexin Ma, Xinge Zhu

机构 * ShanghaiTech University(上海科技大学) Yinwang Intelligent Technology Co. Ltd.(英伟达智能科技有限公司)

专题命中 长视频与时序推理 :video generation(abstract);分类 cs.CV

Comments Accepted for 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 视频数据与评测 3 篇

2312.04817 2025-09-03 cs.CV 79%

LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Hongjie Zhang, Lu Dong, Yi Liu, Yifei Huang, Yali Wang, Limin Wang, Yu Qiao

机构 * OpenGVLab, Shanghai AI Laboratory, China(OpenGVLab,上海人工智能实验室) University of Science and Technology of China(中国科学技术大学) Honor Device Co.,Ltd(荣耀设备有限公司) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院) Nanjing University(南京大学)

专题命中 视频数据与评测 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00484 2025-09-03 cs.CV cs.AI 79%

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding

Zhihong Zhang, Xiaojian Huang, Jin Xu, Zhuodong Luo, Xinzhi Wang, Jiansheng Wei, Xuejin Chen

机构 * University of Science and Technology of China(中国科学技术大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 视频数据与评测 :video understanding(title,abstract);分类 cs.CV

Comments https://videorewardbench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04549 2025-09-03 cs.CV cs.AI cs.MM 73%

MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning

Quang-Trung Truong, Yuk-Kwan Wong, Vo Hoang Kim Tuyen Dang, Rinaldi Gotama, Duc Thanh Nguyen, Sai-Kit Yeung

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Ho Chi Minh University of Science(胡志明市科学大学) Deakin University(德肯大学)

专题命中 视频数据与评测 :video generation(abstract);video understanding(abstract);分类 cs.CV、cs.MM

Comments Published at ACMMM2025 (Dataset track)

详情

展开后加载摘要…

URL PDF HTML 收藏