C2F-Space: Coarse-to-Fine Space Grounding for Spatial Instructions using Vision-Language Models
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract)
Comments 16 pages, 12 figures
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract)
Comments 16 pages, 12 figures
零样本开放词汇人体运动接地与测试时训练
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
AI总结 本文提出ZOMG框架,通过语言语义分区和软掩码优化,在无需标注或微调的情况下实现零样本开放词汇的人体运动接地,提升了运动理解的性能和效率。
机构 * California Institute of Technology(加州理工学院) ; Cedars-Sinai Medical Center(Cedars-Sinai 医疗中心)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted as proceedings paper for ML4H 2025
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(abstract)