arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-09-16 至 2025-09-16 共收录 11 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 4 篇

2509.12145 2025-09-16 cs.CV 79%

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

Hyolim Kang, Yunsu Park, Youngbeom Yoo, Yeeun Choi, Seon Joo Kim

机构 * Yonsei University(延世大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11360 2025-09-16 cs.CV 57%

GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration

Wan Xu, Feng Zhu, Yihan Zeng, Yuanfan Guo, Ming Liu, Hang Xu, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨理工大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03235 2025-09-16 cs.CV cs.AI 57%

Enhancing Traffic Incident Response through Sub-Second Temporal Localization with HybridMamba

Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

机构 * Department of Computer Science, Iowa State University, Ames, IA, USA(计算机科学系,爱荷华州立大学) Department of Civil, Construction and Environmental Engineering, Iowa State University, Ames, IA, USA(土木、建设与环境工程系,爱荷华州立大学)

专题命中 视频理解 :video-language(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01020 2025-09-16 cs.CV 57%

Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation

Junyu Xie, Tengda Han, Max Bain, Arsha Nagrani, Eshika Khandelwal, Gül Varol, Weidi Xie, Andrew Zisserman

机构 * Visual Geometry Group, University of Oxford(牛津大学视觉几何组) CVIT, IIIT Hyderabad(海得拉巴印度理工学院计算机视觉研究所) LIGM, École des Ponts ParisTech(巴黎理工学院路易-狄塞尔数学与计算机科学实验室) SAI, Shanghai Jiao Tong University(上海交通大学人工智能研究所)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments ICCV 2025. Project Page: https://www.robots.ox.ac.uk/vgg/research/shot-by-shot/

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 1 篇

2509.11092 2025-09-16 cs.CV cs.AI 83%

PanoLora: Bridging Perspective and Panoramic Video Generation with LoRA Adaptation

Zeyu Dong, Yuyang Yin, Yuqi Li, Eric Li, Hao-Xiang Guo, Yikai Wang

机构 * School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) Beijing Jiaotong University(北京交通大学) The City College of New York(纽约城市学院) Skywork AI

专题命中 视频生成 :video generation(title,abstract);video diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频问答 3 篇

2509.11796 2025-09-16 cs.CV 79%

FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning

Haodong Chen, Haojian Huang, XinXiang Yin, Dian Shao

机构 * School of Automation, Northwestern Polytechnical University(自动化学院,西北工业大学) The University of Hong Kong(香港大学) School of Software, Northwestern Polytechnical University(软件学院,西北工业大学) Unmanned System Research Institute, Northwestern Polytechnical University(无人系统研究院,西北工业大学)

专题命中 视频问答 :video understanding(title,abstract);分类 cs.CV

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21403 2025-09-16 cs.CV 70%

Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering

Yumeng Shi, Quanyu Long, Wenya Wang

机构 * Nanyang Technological University(南洋理工大学)

专题命中 视频问答 :video language model(abstract);long video(abstract);分类 cs.CV

Comments Accepted to EMNLP 2025 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11862 2025-09-16 cs.CV cs.AI cs.LG 57%

Bridging Vision Language Models and Symbolic Grounding for Video Question Answering

Haodi Ma, Vyom Pathak, Daisy Zhe Wang

机构 * Univerisy of Florida(佛罗里达大学)

专题命中 视频问答 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 长视频与时序推理 1 篇

2504.00879 2025-09-16 cs.CV 57%

GISE-TTT:A Framework for Global InformationSegmentation and Enhancement

Fenglei Hao, Yuliang Yang, Ruiyuan Su, Zhengran Zhao, Yukun Qiao, Mengyu Zhu

专题命中 长视频与时序推理 :long video(abstract);分类 cs.CV

Comments The manuscript requires further improvement

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 视频数据与评测 2 篇

2509.11866 2025-09-16 cs.CV 57%

Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding

Meng Luo, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Li Zheng, Jinxiang Lai, Tianlong Wu, Xinya Du, Jian Li, Siyuan Yan, Jiebo Luo, William Yang Wang, Hao Fei, Mong-Li Lee, Wynne Hsu

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments 25 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11589 2025-09-16 cs.CV 57%

MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment

Yanyun Pu, Kehan Li, Zeyi Huang, Zhijie Zhong, Kaixiang Yang

机构 * Huawei Technologies Co.(华为技术有限公司) South China University of Technology(南方科技大学)

专题命中 视频数据与评测 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏