面向弱监督可供性定位的闭环迁移
Closed-Loop Transfer for Weakly-supervised Affordance Grounding
浏览论文内容
中文总结 AI 辅助
针对弱监督可供性定位中单向知识迁移的局限,本文提出闭环框架 LoopTrans,通过双向迁移、统一跨模态定位与去噪蒸馏提升性能。
中文摘要 AI 辅助
人类只需观察他人与新物体互动,就能完成此前未曾经历的交互。弱监督可供性定位通过使用带有图像级标注的第三人称交互图像,学习在第一人称图像中定位使动作得以发生的物体区域,从而模拟这一过程。然而,仅从第三人称图像中提取可供性知识并将其单向迁移到第一人称图像,限制了以往工作在复杂交互场景中的适用性。为此,本研究提出 LoopTrans,这是一种新颖的闭环框架,不仅将知识从第三人称图像迁移到第一人称图像,还反向迁移以增强第三人称知识提取。在 LoopTrans 中,研究引入了若干创新机制,包括统一跨模态定位和去噪知识蒸馏,以弥合以物体为中心的第一人称图像与以交互为中心的第三人称图像之间的领域差距,同时增强知识迁移。实验表明,LoopTrans 在图像和视频基准的所有指标上均取得持续提升,甚至能够处理物体交互区域被人体完全遮挡的挑战性场景。
英文摘要
Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions that enable actions on egocentric images, using exocentric interaction images with image-level annotations. However, extracting affordance knowledge solely from exocentric images and transferring it one-way to egocentric images limits the applicability of previous works in complex interaction scenarios. Instead, this study introduces LoopTrans, a novel closed-loop framework that not only transfers knowledge from exocentric to egocentric but also transfers back to enhance exocentric knowledge extraction. Within LoopTrans, several innovative mechanisms are introduced, including unified cross-modal localization and denoising knowledge distillation, to bridge domain gaps between object-centered egocentric and interaction-centered exocentric images while enhancing knowledge transfer. Experiments show that LoopTrans achieves consistent improvements across all metrics on image and video benchmarks, even handling challenging scenarios where object interaction regions are fully occluded by the human body.
发表机构
- ShanghaiTech University(上海科技大学)
- School of Computer Science and Engineering(计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。