Bi-MoDe:基于双边控制与修饰符条件解码的模仿学习,用于执行速度与接触强度的调制
Bi-MoDe: Bilateral Control-based Imitation Learning via Modifier-Conditioned Decoding for Modulation of Execution Speed and Contact Intensity
浏览论文内容
中文总结 AI 辅助
Bi-MoDe通过修饰符条件解码框架,在双边控制模仿学习中实现执行速度与接触强度的灵活调制,并在白板擦拭任务上验证了其有效性。
中文摘要 AI 辅助
基于双边控制的模仿学习同时捕捉位置和力信息,因此非常适合接触丰富的操作任务。然而,现有方法为操作员提供的在推理时指定学习任务应如何执行的手段有限,例如缓慢或快速、轻柔或用力。我们提出了Bi-MoDe,一种修饰符条件解码框架,通过adaLN-Zero将受约束的潜在向量注入Transformer动作解码器的每一层,使行为指令能够直接影响动作块生成。我们在真实世界的白板擦拭任务上评估了该方法,并结合了时间与物理修饰符的组合。Bi-MoDe在物理指令遵循方面优于动作分块基线,同时保持相当的时间控制能力。消融研究进一步表明,解码器条件化与潜在空间组合相互作用,且它们的组合对于准确的物理指令遵循至关重要。附加材料可在以下网址获取:this https URL
英文摘要
Bilateral control-based imitation learning captures both position and force information, making it well suited to contact-rich manipulation. However, existing approaches provide limited means for an operator to specify how a learned task should be executed at inference time, such as slowly or quickly, gently or firmly. We propose Bi-MoDe, a modifier-conditioned decoding framework that injects a constrained latent into every layer of the Transformer action decoder via adaLN-Zero, allowing behavioral directives to directly influence action-chunk generation. We evaluate the method on a real-world whiteboard wiping task with combinations of temporal and physical modifiers. Bi-MoDe improves physical directive following over the action-chunking baseline while maintaining comparable temporal control. An ablation further shows that decoder conditioning and latent-space composition interact, and that their combination is important for accurate physical directive following. Additional material is available at the https://mertcookimg.github.io/bi-mode/
发表机构
- The University of Osaka(大阪大学)
- Kobe University(神户大学)
机构由 AI 辅助整理,请以论文原文为准。