arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TLC-DiT:任务对齐的局部视觉条件用于鲁棒的多任务机器人操作

TLC-DiT: Task-Aligned Local Visual Conditioning for Robust Multitask Robot Manipulation

Xianbo Cai, Hideyuki Ichiwara, Zihang Wang, Yijun Lu, Tetsuya Ogata

arXiv 2609.34297首次发表:更新:

发表机构

Waseda University; SB Intuitions Corp.; National Institute of Advanced Industrial Science and Technology (AIST)(早稻田大学; SB Intuitions 公司; 日本产业技术综合研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TLC-DiT通过任务对齐的局部视觉条件扩展多任务扩散策略,在LIBERO上达到93.5%成功率,显著提升鲁棒性并支持可视化检查。

AI 中文摘要

语言条件下的机器人策略在多任务操作方面取得了明显进展,但任务相关的局部视觉证据通常隐藏在视觉骨干网络或注意力层内部。这导致策略难以检查,并且在视觉变化下表现脆弱,这是缺乏显式、任务对齐的局部视觉通道的两个症状。我们提出了TLC-DiT,这是多任务扩散变换器(DiT)策略的即插即用扩展,它增加了显式的任务引导局部视觉特征图,而不改变扩散目标或动作生成过程。对于每个相机视图,冻结的DINOv2补丁特征通过CLIP任务嵌入的FiLM进行调制,并由轻量级CoordConv CNN适配器细化为平滑的空间图,这些图与原始全局图像、语言、关节状态和时间步条件拼接在一起。在LIBERO上,TLC-DiT达到了93.5%的平均成功率,而多任务DiT为86.5%,SmolVLA为79.25%。在LIBERO-plus上,总成功率从54.07%提高到57.24%,在相机、背景和传感器噪声变化下取得了更大的提升。在真实世界的双臂任务中,TLC-DiT将茶包放置完成率从44%提高到89%,同时保持了相当的开箱匹配性能。特征图可视化证实,模型在多个视图和扰动下关注任务相关区域,提供了一种直接检查视觉证据的方法。

英文摘要

Language-conditioned robot policies have made clear progress in multitask manipulation, but task-relevant local visual evidence usually stays hidden inside a visual backbone or attention layers. This leaves the policy difficult to inspect and fragile under visual change, two symptoms of a missing explicit, task-aligned local visual channel. We present TLC-DiT, a plug-in extension of the Multitask Diffusion Transformer (DiT) policy that adds explicit task-guided local visual feature maps without changing the diffusion objective or the action-generation process. For each camera view, frozen DINOv2 patch features are modulated by the CLIP task embedding through FiLM and refined by a lightweight CoordConv CNN adapter into smooth spatial maps, which are concatenated with the original global image, language, joint-state, and timestep conditions. On LIBERO, TLC-DiT reaches a 93.5% average success rate, compared with 86.5% for Multitask DiT and 79.25% for SmolVLA. On LIBERO-plus, the total success rate improves from 54.07% to 57.24%, with larger gains under camera, background, and sensor-noise changes. In real-world bimanual tasks, TLC-DiT raises Teabag Putting completion from 44% to 89% while maintaining comparable Match Box Opening performance. Feature-map visualizations confirm that the model attends to task-relevant regions across views and perturbations, providing a direct way to inspect the visual evidence.

Comments6 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑