Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
Journal ref NeurIPS2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
Journal ref NeurIPS2025
机构 * Harbin University of Science(哈尔滨理工大学) ; Technology The University of Melbourne(技术墨尔本大学)
专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 eess.AS
机构 * Advanced Research Center on Electronic System (ARCES)(电子系统先进研究中心) ; Department of Computer Science and Engineering (DISI)(计算机科学与工程系) ; University of Bologna(博洛尼亚大学)
专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV
Comments ICCV 2025. Code: https://github.com/bartn8/depthanyevent/ Project Page: https://bartn8.github.io/depthanyevent/
专题命中 视频多模态 :multimodal(abstract)
Comments submitted invited review for the `Identifiability, estimation and uncertainty in mathematical modelling' issue in Current Opinion in Systems Biology
Journal ref Curr. Opin. Syst. Biol. 42, 100555 (2025)