arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-22 至 2025-10-22 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 5 篇

2510.17305 2025-10-22 cs.CV cs.MM 84%

LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding

ZhaoYang Han, Qihan Lin, Hao Liang, Bowen Chen, Zhou Liu, Wentao Zhang

机构 * Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学)

专题命中 视频多模态 :omni-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments Submitted to ARR Rolling Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18411 2025-10-22 cs.CL cs.LG 79%

DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding

Yue Jiang, Jichu Li, Yang Liu, Dingkang Yang, Feng Zhou, Quyu Kong

机构 * Fudan University(复旦大学) Center for Applied Statistics and School of Statistics, Renmin University of China(应用统计中心和中国人民大学统计学院) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算高级创新中心)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18726 2025-10-22 cs.CV 57%

IF-VidCap: Can Video Caption Models Follow Instructions?

Shihao Li, Yuanxing Zhang, Jiangtao Wu, Zhide Lei, Yiwen He, Runzhe Wen, Chenxi Liao, Chengkang Jiang, An Ping, Shuo Gao, Suhan Wang, Zhaozhou Bian, Zijun Zhou, Jingyi Xie, Jiayi Zhou, Jing Wang, Yifan Yao, Weihao Xie, Yingshui Tan, Yanghai Wang, Qianqian Xie, Zhaoxiang Zhang, Jiaheng Liu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments https://github.com/NJU-LINK/IF-VidCap

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02175 2025-10-22 cs.RO cs.CV cs.LG 57%

VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching

Siyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu, Tao Huang, Chang Xu

机构 * School of Computer Science, University of Sydney(悉尼大学计算机科学学院) John Hopcropt Center for Computer Science, Shanghai Jiao Tong University(上海交通大学约翰·霍普克罗夫特计算机科学中心)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22205 2025-10-22 cs.RO 50%

From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment

Ke Ye, Jiaming Zhou, Yuanfeng Qiu, Jiayi Liu, Shihui Zhou, Kun-Yu Lin, Junwei Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The University of Hong Kong(香港大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视频多模态 :multimodal(abstract)

Comments More details and videos can be found at: https://yipko.com/super-mimic

详情

展开后加载摘要…

URL PDF HTML 收藏