arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36605cs.RO

RoboChrono:用于流式任务理解的真实机器人基准

RoboChrono: A Real Robot Benchmark for Streaming Task Understanding

Yuzhou Wu, Longteng Fan, Zimeng Li, Yu Wanchan, Ting Zhang, Yiyang Ma, Shihao Li, Wei Ying, Jianbin Qin, Jiajian Jing, Fangwen Chen, Yifan Wu, Zichen Zhang, Rui… 展开作者

Yuzhou Wu, Longteng Fan, Zimeng Li, Yu Wanchan, Ting Zhang, Yiyang Ma, Shihao Li, Wei Ying, Jianbin Qin, Jiajian Jing, Fangwen Chen, Yifan Wu, Zichen Zhang, Ruiqi Yang, Weibin Kong, Yihang Xu, Haoran Liu, Zonghang He, Xuyang Liu, YiFan Xiong, Siteng Huang, Tao Xu, Zhuo Xu, Long Chen, Ruoxiang Li

首次发表
浏览论文内容

中文总结 AI 辅助

RoboChrono是用于流式任务理解的真实机器人基准,含39个场景和34,713个实例,评估18个视觉-语言模型,揭示视觉匹配与时间排序能力不总一致,并表明下一动作预测可依赖任务先验。

中文摘要 AI 辅助

理解正在进行的机器人操作需要模型将视觉观察与交互历史和任务进度相关联。我们引入了RoboChrono,一个用于流式任务理解的基准,包含39个场景和34,713个评估实例,这些实例基于真实机器人执行以及补充的裸手人类录像构建。该基准评估了七项任务,分为识别、对齐和时间定位三类,涵盖动作理解与预测、视觉对应、时间排序和动作定位。对18个视觉-语言模型的零样本评估揭示了任务间的显著差异。GPT-6-Astra在帧匹配上达到98.3%的准确率,但在帧排序上仅为68.3%,而RynnBrain1.1-122B-A10B表现出更大的差距,分别达到95.4%和32.9%。对五个开放权重模型在匹配问题上的输入消融进一步揭示了对视觉证据的不同依赖:移除视觉观察使当前动作识别准确率下降22.1个百分点,而下一动作预测仅下降0.7个百分点。这些发现表明,强大的视觉匹配并不总是与强大的时间排序一致,并提示下一动作预测即使在视觉证据不可用时也能由任务和动作先验支持。RoboChrono为检验这些差异提供了诊断环境,强调了在评估机器人操作中的任务理解时,需要超越总分进行能力特定评估。

英文摘要

Understanding ongoing robot manipulation requires models to interpret visual observations in relation to interaction history and task progress. We introduce RoboChrono, a benchmark for streaming task understanding comprising 39 scenarios and 34,713 evaluation instances, constructed from real robot executions and complementary bare-hand human recordings. The benchmark evaluates seven tasks grouped into recognition, alignment, and temporal grounding, covering action understanding and anticipation, visual correspondence, temporal ordering, and action localization. Zero-shot evaluation of 18 vision-language models reveals substantial differences across tasks. GPT-6-Astra achieves 98.3% accuracy on Frame Matching but 68.3% on Frame Ordering, while RynnBrain1.1-122B-A10B exhibits a larger gap, reaching 95.4% and 32.9%, respectively. Input ablations on matched questions with five open-weight models further reveal distinct dependencies on visual evidence: removing visual observations reduces Current Action Recognition accuracy by 22.1 percentage points, whereas Next Action Prediction decreases by only 0.7 points. These findings show that strong visual matching does not consistently coincide with strong temporal ordering, and suggest that next-action prediction can be supported by task and action priors even when visual evidence is unavailable. RoboChrono provides a diagnostic setting for examining these differences, highlighting the need for capability-specific evaluation beyond aggregate scores when assessing task understanding in robot manipulation.

发表机构

  • Tianji Tec.(天机科技)
  • Shenzhen University(深圳大学)
  • General Intelligence Machine(通用智能机器)
  • Huazhong Agricultural University(华中农业大学)
  • Yanbian University(延边大学)
  • South China Agricultural University(华南农业大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Hong Kong Polytechnic University(香港理工大学)
  • Beijing Jiaotong University(北京交通大学)
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑