CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
CAPE:一种基于CLIP的互补热图线索点集用于具身参照理解
机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) ; Istanbul Technical University(伊斯坦布尔技术大学) ; Carnegie Mellon University(卡内基梅隆大学) ; KIT Campus Transfer GmbH (KCT)(KIT校园转移有限责任公司)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
AI总结 CAPE通过双模型框架和CLIP-aware Pointing Ensemble模块,提升具身参照理解任务中指向线索的多模态推理能力,实现75.0 mAP的高精度表现。
Comments Accepted by WACV 2026