arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

成功即一切?探究输入扰动对桌面操作任务中VLA行为的影响

Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks

Sophie Higham, Riccardo Andrea Izzo, Matteo Matteucci, Alessandro Suglia

arXiv 2610.01351首次发表:更新:

发表机构

University of Edinburgh; Politecnico di Milano(爱丁堡大学; 米兰理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种与基准无关的评估框架,通过扩展LIBERO基准,在多种扰动下度量VLA模型成功轨迹的行为变化,表明仅靠任务成功率不足以全面评估鲁棒性,需结合行为指标。

AI 中文摘要

视觉-语言-动作(VLA)模型在机器人操作任务基准上取得了较高的任务成功率。近期,人们越来越重视评估VLA模型对扰动的鲁棒性。然而,这种鲁棒性仍然主要通过任务成功率(TSR)来衡量。在这项工作中,我们提出了一种与基准无关的评估框架,通过刻画成功轨迹在扰动下如何被执行来度量模型的行为鲁棒性。我们通过扩展广泛使用的LIBERO和LIBERO-Plus基准来实现这一方法。在三种最先进的VLA模型、四个LIBERO任务套件和七种扰动条件下,我们评估了典型成功行为及其变异性的变化,包括运动平滑度、效率和夹爪行为等指标。我们发现扰动可以改变成功轨迹的行为,这一现象不一定能从TSR单独推断出来。在LIBERO套件中,我们识别出在相同扰动条件下,最先进的VLA模型达到相当的TSR,但成功轨迹上的行为却存在显著差异的情况。因此,为了更全面地评估任务性能,我们认为合适的鲁棒性度量不仅应捕获任务是否完成,还应捕获机器人在完成任务过程中的行为方式。在评估VLA模型的鲁棒性时,TSR可以辅以行为评估指标,这些指标刻画机器人成功执行任务的性质和变异性。

英文摘要

Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation task benchmarks. More recently, there has been an emphasis on evaluating the robustness of VLA models to perturbations. However, this robustness is still predominantly measured through Task Success Rate (TSR). In this work, we propose a benchmark-agnostic evaluation framework to measure the behavioural robustness of models by characterising how successful trajectories are executed under perturbation. We implement this methodology by extending the widely-used LIBERO and LIBERO-Plus benchmarks. Across three state-of-the-art VLA models, four LIBERO task suites and seven perturbation conditions, we evaluate changes in both typical successful behaviour and its variability, including metrics of motion smoothness, efficiency and gripper behaviour. We find that perturbations can alter the behaviour of successful trajectories, a phenomenon which cannot necessarily be inferred from TSR alone. Across LIBERO suites, we identify cases where state-of-the-art VLA models achieve comparable TSR under the same perturbation condition, yet behaviour on successful trajectories diverges substantially. Therefore, to have a more robust assessment of task performance, we argue that suitable measures of robustness should capture not only whether a task is completed, but also how the robot behaves while completing it. When evaluating the robustness of VLA models, TSR may be complemented by behavioural evaluation metrics that characterise the nature and variability of successful task execution by robots.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑