arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21659cs.ROcs.AI

基于结果条件化的末端执行器几何在视觉-语言-动作策略中的应用

Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies

  • East China Normal University(华东师范大学)
  • Dalian Maritime University(大连海事大学)

机构由 AI 辅助整理,请以论文原文为准。

Xingyu Lin, Zhuang Li, Zhongrun Wu, Shouquan Zhou, Dehui Du

AI总结:

本研究通过15,000次LIBERO rollout分析四种VLA策略的末端执行器几何,发现双成功策略对距离显著小于单成功对,且该排序在多种表示下稳定,但低距离不意味着可互换性。

AI中文摘要:

视觉-语言-动作(VLA)策略通过不同的动作接口解决相同的操作任务,但任务成功本身并不能确定其物理执行是否一致。我们研究了来自四种策略的15,000次闭环LIBERO rollout中的跨策略末端执行器几何。主要的干净条件分析形成了3,600对配置匹配且因此相互依赖的策略对。双成功对的中位归一化动态时间规整距离为0.0120米,而恰好一个策略成功时为0.0380米。这一排序在每个任务、每个策略对以及九种采样和带限表示中均成立;然而,该比率在不同表示间变化数倍,因此我们报告方向而非固定倍数。双失败对分离更远,但依赖于稀疏且不均匀的支持,因此我们将其作为探索性结果报告。在成功执行中,伙伴替换在不同任务间的分离大于不同初始状态间的分离。匹配的基线仍揭示了可测量的、异质的残余策略差异,因此低跨策略距离并不意味着可互换性。成功执行与同任务演示之间的距离大约相当于这些演示彼此之间的距离,这与任务相关几何一致,但未将训练数据重叠与任务约束分开。一个常见的72动作窗口保持了排序但减小了其幅度;端点和持续时间调整同样相对于双成功对保留了正混合结果系数,尽管其幅度依赖于规格。在复合视觉压力下,策略排名和配对组成同时变化。

英文摘要:

Vision-language-action (VLA) policies solve the same manipulation task through different action interfaces, but task success alone does not establish whether their physical executions agree. We study cross-policy end-effector geometry in 15,000 closed-loop LIBERO rollouts from four policies. The primary clean-condition analysis forms 3,600 configuration-matched, and therefore dependent, policy pairs. Both-success pairs have a median normalized dynamic time warping distance of 0.0120 m versus 0.0380 m when exactly one policy succeeds. This ordering holds in every task, every policy pair, and nine sampling and band-limited representations; however, the ratio varies severalfold across representations, so we report the direction rather than a fixed multiple. Both-failure pairs are more separated again but rest on thin, uneven support, so we report them as exploratory. Within successful executions, partner replacements separate more across tasks than across initial states. A matched baseline still reveals measurable, heterogeneous residual policy differences, so a low cross-policy distance does not imply interchangeability. Successful executions sit about as far from same-task demonstrations as those demonstrations sit from each other, compatible with task-associated geometry without separating training-data overlap from task constraints. A common 72-action window preserves the ordering but reduces its magnitude; endpoint and duration adjustment likewise leaves a positive mixed-outcome coefficient relative to both-success pairs, though its magnitude is specification-dependent. Under composite visual stress, policy rankings and pair composition change together.

补充信息

↑