arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-01-13 至 2026-01-13 共收录 3 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 3 篇

2601.06309 2026-01-13 cs.CV cs.AI 83%

VideoWeave: A Data-Centric Approach for Efficient Video Understanding

VideoWeave:一种以数据为中心的高效视频理解方法

Zane Durante, Silky Singh, Arpandeep Khatua, Shobhit Agarwal, Reuben Tan, Yong Jae Lee, Jianfeng Gao, Ehsan Adeli, Li Fei-Fei

机构 * Stanford University(斯坦福大学) Microsoft Research(微软研究院) University of Wisconsin - Madison(威斯康星大学麦迪逊分校)

专题命中 视频理解 :video understanding(title);video-language(abstract);long video(abstract);分类 cs.CV

AI总结 VideoWeave通过重新组织训练数据提升视频语言模型的数据效率,无需修改模型架构,实现更高准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06843 2026-01-13 cs.CV cs.CL 79%

Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models

边看边说:解锁多模态大语言模型的实时视频理解能力

Junyan Lin, Junlong Tong, Hao Wu, Jialiang Zhang, Jinming Liu, Xin Jin, Xiaoyu Shen

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT) Shanghai Jiao Tong University(上海交通大学) Ocean University of China(中国海洋大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 本文提出并行流式框架,通过三种设计解决多模态大语言模型在实时视频理解中的位置连续性约束问题,实现边看边说的实时交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20876 2026-01-13 cs.RO 50%

Proprioception Enhances Vision Language Model in Generating Captions and Subtask Segmentations for Robot Task

本体感知增强视觉语言模型在为机器人任务生成描述和子任务分割中的应用

Kanata Suzuki, Shota Shimizu, Tetsuya Ogata

机构 * Faculty of Science and Engineering, Waseda University(工学部,早稻田大学) Artificial Intelligence Laboratory, Fujitsu Limited(Fujitsu 人工智能实验室) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究院)

专题命中 视频理解 :video understanding(abstract)

AI总结 本研究通过引入本体感知数据,提升视觉语言模型在机器人任务描述和子任务分割中的性能,以增强机器人模仿学习效率。

详情

展开后加载摘要…

URL PDF HTML 收藏