arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-01-21 至 2026-01-21 共收录 2 信号源:cs.CV, eess.IV, cs.MM

1. 视频数据与评测 2 篇

2601.13974 2026-01-21 cs.CV 57%

STEC: A Reference-Free Spatio-Temporal Entropy Coverage Metric for Evaluating Sampled Video Frames

STEC:一种无参考的时空熵覆盖度量,用于评估采样视频帧

Shih-Yao Lin

机构 * Independent Researcher(独立研究者)

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

AI总结 STEC是一种无参考的时空熵覆盖度量,用于评估视频帧采样的有效性,通过联合建模空间信息强度、时间分散性和非冗余性,提供一种轻量且原则性的采样质量度量。

Comments This paper corresponds to the camera-ready version of a WACV 2026 Workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10054 2026-01-21 q-bio.OT 50%

SurgPub-Video: A Comprehensive Surgical Video Dataset for Enhanced Surgical Intelligence in Vision-Language Model

SurgPub-Video: 一个全面的外科视频数据集,用于增强视觉-语言模型中的外科智能

Yaoqian Li, Xikai Yang, Dunyuan Xu, Yang Yu, Litao Zhao, Xiaowei Hu, Jinpeng Li, Pheng-Ann Heng

专题命中 视频数据与评测 :video understanding(abstract)

AI总结 SurgPub-Video数据集和SurgLLaVA-Video模型通过提供高质量外科视频数据和专门的视觉-语言模型,提升了手术场景分析的智能水平。

Journal ref AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏