Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
专题命中 视觉定位与Grounding :grounding(title,abstract);LLaVA(abstract);multimodal large language model(abstract);MLLM(abstract)
Journal ref NeurIPS2025
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :grounding(title,abstract);LLaVA(abstract);multimodal large language model(abstract);MLLM(abstract)
Journal ref NeurIPS2025
机构 * Syracuse University(苏塞克斯大学) ; Amazon(亚马逊)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments This is the pre-peer-review version of a journal paper; the repo is available at: https://github.com/pnnl/prompt2control