发表机构
Humanoid Robot (Shanghai) Co., Ltd.; E-surfing Digital Life Technology Co., Ltd., China Telecom(人形机器人(上海)有限公司; 中国电信天翼数字生活科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LM-X是一种面向通用机器人操控的可解释动作建模方法,通过多尺度预测信号提升操控成功率,在RoboTwin2.0等任务上显著优于现有模型。
AI 中文摘要
通用视觉-语言-动作(VLA)策略主要通过短程动作预测学习长程行为,但仅能输出采样得到的指令,几乎无法提供额外信息。这造成两个耦合瓶颈:单一动作目标必须隐含地包含任务进度、中间意图与局部可靠性,而这些控制状态在执行过程中始终处于隐藏状态。受生物感知运动控制功能原理启发,我们提出LM-X,其在任务、事件与运动尺度上组织预测,不追求与生物结构的对应关系。模型在线生成三个显式监督信号,直接调控动作生成:待完成回报(RTG)衡量可见任务进度,待完成事件(ETG)识别下一个语义转换,异方差动作流通过传播方差估计局部可靠性。因此,解释是控制的固有属性,而非事后生成。在64块NVIDIA B200 GPU上进行耗时20天的预训练前,我们通过受控的五任务预训练门控验证了设计:完整模型的成功率仅动作骨干网络提升16.0个百分点,仅最强单头变体提升10.8个百分点。随后,我们在超过20000小时的真实机器人轨迹(含超过1000小时的失败策略 rollout)上训练LM-X。LM-X在50个随机困难RoboTwin2.0任务上的平均成功率达74.1%,而GR00T N1.7为55.4%;在7个真实机器人任务上为68.6%,GR00T N1.7为50.7%。RTG可追踪语义进度与可见退化,方差在犹豫与振荡控制期间上升。这些结果表明,显式多时间尺度预测状态可强化控制,同时提供可解释的内部估计。
英文摘要
Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: actions are exposed, but their explanatory state is not. They provide no native account of three explanatory signals: task progress, the next semantic transition, or local command reliability. Prior work shows that progress and event structure aid long-horizon control and that uncertainty supports monitoring; however, such capabilities are typically added or extracted only after action pretraining. The field therefore lacks a VLA foundation model whose explanatory state is jointly pretrained with control. Drawing on biological sensorimotor organization, in which outcome-sensitive, event-segmented, and probabilistic predictions structure behavior, we introduce LM-X. LM-X learns three directly supervised online signals: return-to-go (RTG) estimates visible progress and state quality; event-to-go (ETG) predicts the action sequence to the next semantic event; and heteroscedastic action-flow variance reports local command reliability. RTG conditions ETG and both condition action generation; uncertainty is estimated inside the action expert, making explanation part of control rather than a post-hoc description. We pretrain LM-X on more than 20,000 hours of heterogeneous real-robot trajectories, including over 1,000 hours of failed rollouts. LM-X achieves 74.1\% success on 50 randomized-hard RoboTwin2.0 tasks and 73.5\% on seven real-robot tasks, compared with 55.4\% and 50.7\% for GR00T N1.7. Its signals track progress and regression, anticipate event-scale motion, detect high-error actions, and provide advance failure warning. These results establish LM-X as an explainable VLA foundation model that couples transparent predictive state with stronger generalist control. Github: https://github.com/loongOpen/LoongWu-LM-X-VLA