arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Behavior-Skill:用于评估长 horizon 任务中视觉-语言-动作策略的细粒度基准

Behavior-Skill: A Fine-Grained Benchmark for Evaluating Vision-Language-Action Policies in Long-Horizon Tasks

Chunyun Ma, Lun Luo, Xingjian Luo, Xiexing Feng, Hang Zhang, Wei Liu, Feng Qiao, Yaonan Wang, Huimin Lu, Xieyuanli Chen

arXiv 2608.30536首次发表:更新:

发表机构

XPeng Inc.; The Chinese University of Hong Kong; Hunan University(小鹏汽车; 香港中文大学; 湖南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出细粒度基准 Behavior-Skill,含 235492 个技能实例,引入轨迹与技能级指标,实验发现接触丰富操作技能是长 horizon VLA 策略瓶颈,可用于分析改进该类策略。

AI 中文摘要

长 horizon 移动操作任务的可靠执行仍具挑战性,因为整体任务成功取决于多个组成技能的成功完成。然而,现有基准仍主要依赖全任务 rollout 及聚合的任务级指标,导致中间阶段的失败难以观察和分析。我们提出 Behavior-Skill,一个围绕可执行组成技能重新构建长 horizon 任务学习与评估的基准。该基准包含来自 50 个家庭任务、34 种语义技能类别的 10,000 次演示中的 235,492 个技能实例,每个实例将技能指令与对齐的观测-动作片段配对,并关联可恢复的中间状态和技能成功条件,以支持有效前置条件下的独立评估。我们进一步引入轨迹级和技能级指标,以表征超越聚合任务成功的策略能力。在完整 50 任务基准上对包括 pi0.5 和 GR00T 在内的代表性 VLA 策略开展的大量实验显示,失败在技能间高度非均匀,接触丰富的操作技能构成持续瓶颈。这些结果表明,Behavior-Skill 通过暴露中间能力轮廓,补充了全任务评估,用于分析和改进长 horizon VLA 策略。Behavior-Skill 可在该 https URL 公开获取。

英文摘要

Reliable execution of long-horizon mobile manipulation tasks remains challenging because overall task success depends on the successful completion of multiple constituent skills. Existing benchmarks, however, still rely primarily on full-task rollouts and aggregate task-level metrics, making intermediate failures difficult to observe and analyze. We present Behavior-Skill, a benchmark that reformulates the learning and evaluation of long-horizon tasks around executable constituent skills. It contains 235,492 skill instances from 10,000 demonstrations across 50 household tasks and 34 semantic skill categories. Each instance pairs a skill instruction with an aligned observation-action segment, and is further associated with a restorable intermediate state and a skill success condition to enable independent evaluation under valid preconditions. We further introduce trajectory-level and skill-level metrics to characterize policy capability beyond aggregate task success. Extensive experiments across representative VLA policies including pi0.5 and GR00T on the complete 50-task benchmark show that failures are highly non-uniform across skills, with contact-rich manipulation skills forming persistent bottlenecks. These results demonstrate that Behavior-Skill complements full-task evaluation by exposing intermediate capability profiles for analyzing and improving long-horizon VLA policies. Behavior-Skill is publicly available at https://github.com/nubot-nudt/Behavior-Skill.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑