arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboSPA:VLA模型能否超越简单场景与短视任务?

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang

arXiv 2609.05324首次发表:更新:

发表机构

Zhejiang University; University of Electronic Science and Technology of China; South China Normal University(浙江大学; 电子科技大学; 华南师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究推出RoboSPA数据集与基准,评估VLA模型在复杂空间推理、长时程规划等任务的表现,发现当前VLA模型在相关任务仍存在不足,为开发更优具身智能体提供支撑。

AI 中文摘要

视觉-语言-动作(Vision-Language-Action,VLA)模型在语言条件下的机器人操纵任务中已展现出良好进展,但现有数据集与基准主要评估预定义设置下的任务完成情况,难以洞察模型在空间与过程复杂度提升时的推理能力。我们推出RoboSPA(Robot Spatial-Procedural Assessment,机器人空间-过程评估),这是用于诊断VLA模型具身推理能力的大规模机器人操纵数据集与基准。RoboSPA聚焦精细空间推理与长时程过程规划两个核心维度,涵盖10类任务、56项基础任务,每项任务设5个难度等级,共生成280种空间模糊性与过程复杂度递增的变体;我们采集了多类具身系统与多样场景下的52.7万条轨迹。除二元成功率外,RoboSPA还引入诊断指标以实现更细致的评估。对代表性VLA模型的实验显示,当前系统仍难以应对复杂空间关系、精确低层执行及高内存需求的规划任务。这些结果表明RoboSPA可作为极具挑战性的诊断基准,用于开发更强大、可靠且泛化能力更强的具身智能体,相关数据与代码可在指定URL获取。

英文摘要

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce \textbf{RoboSPA} (\textbf{Robo}t \textbf{S}patial-\textbf{P}rocedural \textbf{A}ssessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. \texttt{RoboSPA} focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, \texttt{RoboSPA} introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish \texttt{RoboSPA} as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.

CommentsAccepted at the EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑