arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11769cs.RO

人形双臂操纵中的策略诱导手部先验:诊断与缓解初始位姿依赖

Policy-Induced Hand Priors in Humanoid Dual-Arm Manipulation: Diagnosing and Mitigating Initial-Pose Dependence

  • Center for Humanoid Research, Korea Institute of Science and Technology (KIST)(韩国科学技术研究院类人机器人研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Chaeyeon Jung, Juyoun Park

AI总结:

本研究针对人形双臂操纵中VLA策略的初始位姿依赖问题,量化策略诱导手部先验,发现扩大训练数据初始位姿覆盖可提升稳健性,为缓解该依赖提供方法。

AI中文摘要:

视觉-语言-动作(VLA)策略应能在机器人初始构型变化时稳健运行,但总体任务成功率可能掩盖特定位姿下的失败及不当手部选择。本研究调查基于VLA的人形双臂操纵中的初始位姿依赖,将依赖初始条件的早期手部偏好表征为策略诱导手部先验,并使用HandPriorScore(手部先验得分)、残差手部偏差及目标响应性对其量化。对多种策略和17种初始构型的评估显示,存在强烈的初始位姿-策略交互:同一姿态在不同策略间产生显著不同的成功率,而单一策略在不同姿态间表现出巨大性能差异。特定初始手臂构型可抑制或诱导非对称手部偏好,且该效应的方向和强度因策略而异;腕部相机观测也会影响手部选择和任务性能。在训练数据集中扩大初始位姿覆盖范围可显著提升稳健性,而针对低性能构型的定向增强可提高其成功率。对训练构型的比较表明,充分接触目标模拟任务是有益的,而真实或辅助数据的效果取决于位姿覆盖、模拟比例和观测可用性。这些发现表征了位姿条件下的手部先验,确定局部初始手臂构型是手部选择行为的因果控制因素,并证明数据覆盖和训练组成如何影响初始位姿稳健性。

英文摘要:

Vision-language-action (VLA) policies are expected to operate robustly across variations in the robot's initial configuration, yet aggregate task success can conceal pose-specific failures and inappropriate hand selection. This work investigates initial-pose dependence in VLA-based humanoid dual-arm manipulation. We characterize the initial-condition-dependent early hand preference as a policy-induced hand prior and quantify it using HandPriorScore, residual hand bias, and target responsiveness. Evaluations across multiple policies and 17 initial configurations reveal strong initial-pose--policy interactions: the same pose produces substantially different success rates across policies, while a single policy exhibits large performance variation across poses. Specific initial arm configurations can suppress or induce an asymmetric hand preference, with the resulting effect varying in direction and strength across policies. Wrist-camera observations also influence hand selection and task performance. Expanding initial-pose coverage in the training dataset substantially improves robustness, while targeted augmentation around a low-performing configuration increases its success rate. Comparisons across training configurations show that sufficient exposure to the target simulation task is beneficial, whereas the effect of real or auxiliary data depends on pose coverage, simulation ratio, and observation availability. These findings characterize a pose-conditioned hand prior, identify a localized initial arm configuration as a causal handle on hand-selection behavior, and demonstrate how data coverage and training composition affect initial-pose robustness.

↑