arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-11-18 至 2025-11-18 共收录 4 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 4 篇

2511.13054 2025-11-18 cs.CV 84%

ViSS-R1: Self-Supervised Reinforcement Video Reasoning

Bo Fang, Yuxin Song, Qiangqiang Wu, Haoyuan Sun, Wenhao Wu, Antoni B. Chan

机构 * City University of Hong Kong(香港城市大学) Baidu Inc.(百度公司) Tsinghua University(清华大学) The University of Sydney(悉尼大学)

专题命中 视频理解 :video reasoning(title,abstract);video understanding(abstract);分类 cs.CV

Comments Our paper was initially titled "Video-SSR1: Self-Supervised Reinforcement Video Reasoning." Upon noticing its close resemblance to the title of a recently released paper, we have decided to rename our work as "ViSS-R1."

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13644 2025-11-18 cs.CV 79%

CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding

Shrenik Patel, Daivik Patel

机构 * Rutgers University(罗切斯特大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12530 2025-11-18 cs.CV 79%

ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding

Yuan Zhou, Litao Hua, Shilong Jin, Wentao Huang, Haoran Duan

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments Accepted to AAAI 2026. Code is available at: https://github.com/robin-hlt/AAAI26-ReaSon

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23447 2025-11-18 cs.CV 57%

CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition

Jongseo Lee, Joohyun Chang, Dongho Lee, Jinwoo Choi

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Our paper has been accepted to IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏