用于自主可变形操作的TCAM:WBCD 2026赛道4的RMC2冠军系统
TCAM for Autonomous Deformable Manipulation: The RMC2 Champion System for WBCD 2026 Track 4
浏览论文内容
中文总结 AI 辅助
该研究围绕TCAM框架构建自主可变形操作系统,在WBCD 2026赛道4挑战赛中以平均23秒/件的速度完成25件T恤操作,22件达标,获赛道冠军。
中文摘要 AI 辅助
本技术报告介绍了RMC2团队针对WBCD 2026赛道4(可变形操作挑战赛)的冠军解决方案。该任务要求机器人从堆叠物中取出单件T恤,将其装载到打印托盘上,使衣领与目标区域对齐,并平整打印区域,该序列涉及单层分离、可变形运输、精确放置以及接触丰富的表面调整。比赛强烈鼓励完全自主执行,推动了自主解决方案的开发。我们围绕TCAM(TermiBrain因果动作模型)框架构建了完全自主系统,设计原则是硬件、感知、数据和学习应共同降低策略必须处理的物理交互复杂性。为单层织物分离设计的定制3D打印夹具提高了双臂ARX X5平台上的拾取可靠性。以手腕为中心的四摄像头设置,结合用于任务级上下文的上部鱼眼相机和用于近距离夹具-布料接触观察的下部RGB相机。我们将便携式UMI风格演示与在可部署平台上收集的真实机器人演示相结合,以提供广泛的操作先验和特定于部署的动力学。TCAM将这些组件连接成闭环:分析每个轨迹以识别导致其结果的物理因素,驱动有针对性的数据重新收集和策略微调。该策略从多视图VLA主干输出30步末端执行器delta-pose动作块。在最终比赛中,我们的系统装载了25件T恤,每次尝试平均约23秒,其中22件达到了所需的表面平滑度,确保了赛道4的第一名。
英文摘要
This technical report describes the RMC2 Team's champion solution for the WBCD 2026 Track 4: Deformable Manipulation Challenge. The task requires a robot to pick a single T-shirt from a stack, load it onto a printing pallet, align the collar with a target area, and smooth the printing region, a sequence that involves single-layer separation, deformable transport, precise placement, and contact-rich surface adjustment. The competition strongly incentivizes fully autonomous execution, motivating the development of an autonomous solution. We built a fully autonomous system around the TCAM (TermiBrain Causal Action Model) framework, with the design principle that hardware, perception, data, and learning should jointly reduce the physical interaction complexity the policy must handle. A custom 3D-printed gripper designed for single-layer fabric separation improves picking reliability on a dual-arm ARX X5 platform. A wrist-centric four-camera setup pairs upper fisheye cameras for task-level context with lower RGB cameras for close-range gripper-cloth contact observation. We combine portable UMI-style demonstrations with real-robot demonstrations collected on the deployable platform to provide both broad manipulation priors and deployment-specific dynamics. TCAM ties these components into a closed loop: each trajectory is analyzed to identify the physical factors contributing to its outcome, driving targeted data recollection and policy fine-tuning. The policy outputs 30-step end-effector delta-pose action chunks from a multi-view VLA backbone. In the final competition, our system loaded 25 T-shirts at an average of approximately 23 seconds per attempt, with 22 achieving the required surface smoothness, securing first place in Track 4.
发表机构
- TermiTech(特米科技)
机构由 AI 辅助整理,请以论文原文为准。