arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-07 至 2025-10-07 共收录 5 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 5 篇

2509.24008 2025-10-07 cs.CV cs.AI 79%

FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning

Haonan Ge, Yiwei Wang, Kai-Wei Chang, Hang Wu, Yujun Cai

机构 * University of California, Merced(加州大学默塞德分校) University of California, Los Angeles(加州大学洛杉矶分校) The University of Queensland(昆士兰大学)

专题命中 视频理解 :video reasoning(title);video understanding(abstract);分类 cs.CV

Comments Underreview

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04401 2025-10-07 cs.CV cs.AI 57%

Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting

Xuyang Guo, Zekai Huang, Zhenmei Shi, Zhao Song, Jiahao Zhang

机构 * Guilin University of Electronic Technology(桂林电子科技大学) The Ohio State University(俄亥俄州立大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03955 2025-10-07 cs.CV 57%

Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs

Sameep Vani, Shreyas Jena, Maitreya Patel, Chitta Baral, Somak Aditya, Yezhou Yang

机构 * Arizona State University(亚利桑那州立大学) Indian Institute of Technology, Kharagpur(印度理工学院,克拉格浦)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 17 pages, 9 figures, 6 tables. Presents TimeWarp, a synthetic preference data framework to improve temporal understanding in Video-LLMs, showing consistent gains across seven benchmarks. Includes supplementary material in the Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03706 2025-10-07 cs.RO cs.AI cs.CV cs.LG 57%

EmbodiSwap for Zero-Shot Robot Imitation Learning

Eadom Dessalene, Pavan Mantripragada, Michael Maynord, Yiannis Aloimonos

机构 * department of Computer Science, University of Maryland, College Park, MD, 20742(计算机科学系,马里兰大学, College Park, MD)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Video link: https://drive.google.com/file/d/1UccngwgPqUwPMhBja7JrXfZoTquCx_Qe/view?usp=sharing

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11336 2025-10-07 cs.CV 57%

UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks

Peiran Wu, Yunze Liu, Zhengdong Zhu, Enmin Zhou, Junxiao Shen

机构 * University of Bristol(布里斯托大学) Memories.ai Research(Memories.ai研究)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏