arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-11-11 至 2025-11-11 共收录 3 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 3 篇

2507.17047 2025-11-11 cs.CV cs.AI 79%

Controllable Hybrid Captioner for Improved Long-form Video Understanding

Kuleen Sasse, Efsun Sarioglu Kayi, Arun Reddy

机构 * Johns Hopkins University Applied Physics Laboratory(约翰霍普金斯大学应用物理实验室)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18094 2025-11-11 cs.CV cs.AI 57%

UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning

Ye Liu, Zongyang Ma, Junfu Pu, Zhongang Qi, Yang Wu, Ying Shan, Chang Wen Chen

机构 * The Hong Kong Polytechnic University(香港理工大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) vivo Mobile Communication Co.(vivo移动通信公司) MindWingman Technology (Shenzhen) Co., Ltd.(深圳MindWingman技术有限公司)

专题命中 视频理解 :video-language(abstract);分类 cs.CV

Comments NeurIPS 2025 Camera Ready. Project Page: https://polyu-chenlab.github.io/unipixel/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05833 2025-11-11 cs.CV 57%

TYrPPG: Uncomplicated and Enhanced Learning Capability rPPG for Remote Heart Rate Estimation

Taixi Chen, Yiu-ming Cheung

机构 * School of Computing, Binghamton University(计算学院,宾夕法尼亚州立大学) Department of Computer Science, Hong Kong Baptist University(计算机科学系,香港 Baptist 大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments The 6th International Workshop on AI for Social Good in the Connected World (AI4SG)@ IEEE WI-IAT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏