发表机构
University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出技能条件化视觉触觉表征,通过进度引导的稀疏事件记忆,在接触密集型任务中显著降低滑移检测、扭转完成延迟及进度误差,揭示观察形成是多模态表征的核心挑战。
AI 中文摘要
机器人操作整合了视觉、触觉和语言,其重要性在不同阶段会发生变化:视觉引导到达,而触觉则通过其随时间的演变来决定抓取、对齐和接触。然而,现有的多模态操作策略通常使用固定的时间上下文和融合策略,尽管在不同技能中每种模态的贡献会有所变化。我们研究了在原始技能层面上视觉和触觉应如何结合,询问每种技能需要从每个传感器获得什么,并提出了一种技能条件化表征,其中查询的技能条件化融合了特定模态的短期观察令牌,同时关注一个稀疏事件记忆,该记忆保留最近$K$个已执行技能的终端观察。通过在三个接触密集型任务上的技能进度评估,与微调的最先进进度模型相比,它将滑移检测延迟降低了87%,与仅视觉消融相比,将扭转完成延迟降低了67.5%,并通过稀疏事件记忆将盲搜索任务上的进度误差降低了92%。增益恰好集中在由接触或任务历史定义完成的区域。更广泛地说,我们的结果表明,观察形成不仅是策略架构,也是多模态表征中的核心挑战。项目网站:此HTTP URL
英文摘要
Robotic manipulation integrates vision, touch, and language, whose importance shifts across stages: vision guides reaching, while touch, through its evolution over time, decides grasping, alignment, and contact. Yet existing multi-modal manipulation policies typically use fixed temporal contexts and fusion strategies, despite shifts in what each modality contributes across different skills. We study how vision and touch should be combined at the level of primitive skills, asking what each skill needs from each sensor, and propose a skill-conditioned representation in which the queried skill conditions fusion over modality-specific short-term observation tokens while attending to a sparse event memory that retains terminal observations from the last $K$ executed skills. Evaluated by skill progress estimation on three contact-rich tasks, it reduces slip-detection delay by 87% against fine-tuned SOTA progress models, twist-completion delay by 67.5% against a vision-only ablation, and progress error on a blind search task by 92% through sparse event memory. Gains concentrate exactly where completion is defined by contact or task history. More broadly, our results suggest that observation formation not only policy architecture is a central challenge in multi-modal representation. Project Website: http://what-to-attend-what-to-keep.github.io/