arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16556cs.AI

DeepInsight II:从基准测试到机器人的单条轨迹

DeepInsight II: One Trace from Benchmark to Robot

Siyi Li, Yuchen Kang, Wuliang Wang, Zhengjie Zhang, Jiangpin Liu, Jianhao Yao, Jie Chen

首次发表
浏览论文内容

中文总结 AI 辅助

DeepInsight II量化具身层,复现基准结果、通过MotionBench实现仿真到真实机器人的原生迁移,并扩展轨迹定位至带修复动作的交接标签,提供从基准到机器人的经验连续性。

中文摘要 AI 辅助

在物理AI技术栈中,评估成熟度与部署风险呈负相关:基础模型拥有成熟、标准化的评估框架,而实际部署依赖的具身层却因基准特定的模拟器、实体和接口而碎片化。首个DeepInsight报告(v1)通过任务、资源、结果三个抽象概念统一了该技术栈的评估,但其定量证据集中在基础模型层;导航与操作(系统1)、全身控制(系统0)仍为仿真案例研究,物理执行未纳入其实验范围。DeepInsight II保留该底层框架并量化具身部分:首先,在两个导航和四个操作基准的原生协议下复现已发布检查点的参考结果;其次,MotionBench将四个已发布的全身控制器置于同一工作负载与指标约定下,再将合格的同系列群体从并行仿真迁移至匹配的真实机器人试验,其中仿真与物理rollout共享父轨迹标识,同时保留执行域特定记录,使仿真到真实的迁移成为原生缩减而非跨工具链的协调;第三,组合的系统2-1-0研究将轨迹定位扩展为五个基于证据的交接标签,每个标签对应具体的修复动作,具备可测量的可修复性标准,且物理试验在硬件可观测状态下测试相同归因。因此,其贡献并非新的评估架构,而是从基准执行到匹配机器人证据及面向修复的诊断的经验连续性。

英文摘要

Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized harnesses, while the embodied layers on which deployment actually turns remain fragmented across benchmark-specific simulators, embodiments, and interfaces. The first DeepInsight report (v1) unified evaluation across this stack behind three abstractions---task, resource, and result---but its quantitative evidence centered on the foundation-model layer; navigation and manipulation (System 1) and whole-body control (System 0) remained simulation case studies, and physical execution was outside its empirical scope. DeepInsight II keeps that substrate fixed and quantifies the embodied half. First, it reproduces released-checkpoint references across two navigation and four manipulation benchmarks under their native protocols. Second, MotionBench places four released whole-body controllers under one workload and metric contract, then carries a qualified within-family cohort from parallel simulation to matched real-robot trials in which simulated and physical rollouts share a parent trace identity while retaining execution-domain-specific records, making the sim-to-real gap a native reduction rather than a reconciliation across toolchains. Third, a composed System 2--1--0 study extends trace localization into five evidence-grounded handoff labels, each mapped to a concrete repair action, with a measured repairability criterion and physical episodes testing the same attribution under hardware-observable state. The contribution is therefore not a new evaluation architecture, but empirical continuity from benchmark execution to matched robot evidence and repair-oriented diagnosis.

↑