文本到3D策略:用于未见规范泛化的细粒度语言-行为对齐
Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization
浏览论文内容
中文总结 AI 辅助
针对文本到3D策略难以泛化到未见细粒度行为规范的问题,提出T3DP框架,通过令牌级双向对齐实现细粒度语言-行为匹配,显著提升多个基准和真实机器人任务的成功率。
中文摘要 AI 辅助
3D视觉运动策略为空间精确操作提供了坚实基础,然而当前的文本到3D策略难以遵循超出演示所覆盖范围的未见细粒度行为规范。我们将这一挑战视为未见规范泛化问题,其中语言指定了行为上重要的变化,如目标位置、位移或关节状态,这些变化在策略训练中不存在。我们发现,预训练语言表示和传统的全局行为-语言对齐能够捕捉粗略的任务语义,但往往模糊了需要不同行为的邻近规范。我们引入了T3DP,一个用于细粒度语言-行为对齐的文本到3D策略框架。T3DP不是将每个指令和演示压缩为单个全局嵌入,而是保留其局部结构,并在语言元素和行为片段之间建立双向的令牌级对应关系。这直接将细微的语言变化锚定到它们所影响的行为组件上,防止紧密相关的规范在表示空间中坍缩。由此产生的规范敏感语言表示条件化了一个基于点云的3D扩散策略,从而在不修改底层策略架构的情况下,实现对未见行为规范更精确的控制。在Meta-World、ManiSkill和RoboTwin上,T3DP相比全局语言-行为对齐,平均未见规范成功率提高了+11.0-14.2个百分点,在所有15个任务族上均有所提升;在真实机器人任务上,它进一步将平均成功率从47.5%提高到65.0%(+17.5个百分点)。表示和动作探针分析表明,细粒度对齐更好地保留了规范几何和动作相关变化,将局部行为锚定与下游控制联系起来。
英文摘要
3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behaviorally significant variations, such as target position, displacement, or articulated state, that are absent from policy training. We find that pretrained language representations and conventional global behavior-language alignment capture coarse task semantics but often blur nearby specifications that require distinct behaviors. We introduce T3DP, a Text-to-3D Policy framework for fine-grained language-behavior alignment. Rather than compressing each instruction and demonstration into a single global embedding, T3DP preserves their local structures and establishes bidirectional token-level correspondence between linguistic elements and behavioral segments. This directly grounds subtle linguistic variations in the behavior components they affect, preventing closely related specifications from collapsing in the representation space. The resulting specification-sensitive language representation conditions a point-cloud-based 3D diffusion policy, enabling more precise control over unseen behavioral specifications without modifying the underlying policy architecture. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0-14.2 points, with gains on all 15 task families; on real-robot tasks, it further raises average success from 47.5% to 65.0% (+17.5 points). Representation and action-probe analyses show that fine-grained alignment better preserves specification geometry and action-relevant variation, linking local behavior grounding to downstream control.
发表机构
- Nanjing University(南京大学)
- Australian National University(澳大利亚国立大学)
- University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。