arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更快更好?基准缺陷与设计局限扭曲视觉-语言-动作加速评估

Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action Acceleration

Qiwei Chen, Kaijun Zhou, Nuohui Shi, Zhiyang Li, Yuxuan Feng, Jinyu Gu

arXiv 2609.37771首次发表:更新:

发表机构

Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究审计七个模拟操作基准,发现22个基准缺陷和4个设计局限导致免训练加速方法成功率虚高;修复后方法排名反转,表明加速增益可能源于基准问题而非任务执行改进。

AI 中文摘要

模拟操作基准是评估视觉-语言-动作(VLA)策略及其为机器人部署降低推理延迟的加速方法的标配工具。在这些基准上,我们观察到一些免训练的加速方法(这些方法近似基线策略的计算)所测得成功率高于基线本身。仅凭成功率无法确定这些增益是源于更好的任务执行还是评估缺陷。因此,我们调查了这些增益背后的两类基准缺陷:缺陷(bug),即实现与预期任务或评估协议不匹配;以及设计局限,即成功标准和模拟设置未能完全捕捉加速对任务执行的影响。从出现异常增益的任务出发,我们通过将物体轨迹与检查器接受区域对比来定位根本原因,并将由此产生的缺陷分类为任务一致性、初始化和可复现性。将这一审计扩展到七个基准,包括RoboTwin、LIBERO-Plus和VLABench,我们识别出22个此类缺陷和4个设计局限。对于后者,我们修正了宽松的成功检查器,纠正了不切实际的物体质量,并添加了一个偏好更平滑动作的运动感知评分。实验表明,修复缺陷可以逆转方法排名,在某个任务上将基线从最后一位提升至第一位。解决设计局限同样可以消除异常增益:在另一个任务上,基线从落后加速方法21个百分点变为领先5个百分点。因此,归因于加速的增益可能是基准的产物,而非更好的任务执行。我们发布缺陷修复和修订后的基准设置,以支持对VLA加速的可信评估。

英文摘要

Simulated manipulation benchmarks are the standard tool for evaluating vision-language-action (VLA) policies and the acceleration methods that reduce their inference latency for on-robot deployment. On these benchmarks, we observe that some training-free acceleration methods, which approximate the baseline policy's computation, achieve higher measured success rates than the baseline itself. Success rates alone cannot establish whether such gains come from better task execution or from evaluation flaws. We therefore investigate two kinds of benchmark flaws behind these gains: bugs, where the implementation does not match the intended task or evaluation protocol, and design limitations, where success criteria and simulation settings do not fully capture how acceleration affects task execution. Starting from tasks with anomalous gains, we localize root causes by plotting object trajectories against checker acceptance regions, and classify the resulting bugs into task consistency, initialization, and reproducibility. Extending this audit to seven benchmarks, including RoboTwin, LIBERO-Plus, and VLABench, we identify 22 bugs of these types and 4 design limitations. For the latter, we revise permissive success checkers, correct unrealistic object masses, and add a motion-aware score that favors smoother actions. Experiments show that bug fixes can reverse method rankings, moving the baseline from last to first on one task. Addressing design limitations can likewise remove anomalous gains: on another task, the baseline moves from 21 percentage points behind an accelerated method to 5 points ahead. Gains attributed to acceleration can therefore be artifacts of the benchmark rather than better task execution. We release our bug fixes and revised benchmark settings to support trustworthy evaluation of VLA acceleration.

CommentsIncludes appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑