arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-16 至 2025-10-16 共收录 3 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 3 篇

2507.05258 2025-10-16 cs.CV cs.LG 70%

Spatio-Temporal LLM: Reasoning about Environments and Actions

Haozhen Zheng, Beitong Tian, Mingyuan Wu, Zhenggang Tang, Klara Nahrstedt, Alex Schwing

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Code and data are available at https://zoezheng126.github.io/STLLM-website/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16836 2025-10-16 cs.CV cs.AI 62%

Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning

Fanrui Zhang, Dian Li, Qiang Zhang, Jun Chen, Gang Liu, Junxiong Lin, Jiahong Yan, Jiawei Liu, Zheng-Jun Zha

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, USTC(脑启发智能感知与认知国家重点实验室,中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) Tencent QQ(腾讯QQ) Fudan University(复旦大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 34 pages, 25 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13019 2025-10-16 physics.optics 50%

Phase Matching of Orbital Angular Momentum in Rare Earth Ion Doped Solid State Systems

Owen R. Wolfe, Joshua Dugre, Grant Kirkland, R. Krishna Mohan

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏