发表机构
KTH Royal Institute of Technology(瑞典皇家理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出FocusPool注意力池化模块,根据本体感觉上下文选择性聚合CNN中间特征,在仿真和真实实验中成功率分别提升36.2%和41.2%,仅训练5.8%编码器参数。
AI 中文摘要
聚焦于空间局部化、与控制相关的视觉线索已被证明可以通过减少对任务无关视觉变化建模的需求来提高视觉运动策略的数据效率。现有方法通常通过输入预处理(例如在RGB图像或点云中裁剪控制中心或物体中心区域)来施加这种聚焦。然而,这种局部特征是否可以直接从常用的卷积神经网络(CNN)编码特征中提取,仍未得到充分探索。在本文中,我们表明中间CNN特征保留了用于控制的局部视觉上下文,但现有的池化方法无法有效聚合这些特征。我们引入了FocusPool,一种注意力池化模块,它根据中间视觉特征与机器人当前本体感觉上下文的相关性,选择性地聚合这些特征。由此产生的池化表示捕获了任务渐进、控制相关的局部信息,并直接用于策略学习。在仿真和真实世界实验中,FocusPool在策略成功率上比池化方法和显式局部聚焦方法分别提高了36.2%和41.2%,且仅训练了编码器参数的5.8%。
英文摘要
Focusing on spatially localized, control-relevant visual cues has been shown to improve data efficiency in visuomotor policies by reducing the need to model task-irrelevant visual variation. Existing methods often impose this focus through input preprocessing, such as cropping control- or object-centric regions in RGB images or point-clouds. However, it remains underexplored whether such localized features can be exposed directly from commonly used convolutional neural network (CNN) encoded features. In this paper, we show that intermediate CNN features preserve localized visual context for control, but existing pooling methods fail to aggregate it effectively. We introduce FocusPool, an attention pooling module that selectively aggregates intermediate visual features according to their relevance to the robot's current proprioceptive context. The resulting pooled representation captures task-progressive, control-relevant local information and is used directly for policy learning. Across simulation and real-world experiments, FocusPool improves policy success rates over pooling and explicit local focus methods by 36.2% and 41.2%, with training only 5.8% of encoder parameters.
CommentsConference on Robot Learning (CoRL), 2026