arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LM-X:面向通用机器人操控的、具备进度、事件与不确定性预测能力的可解释动作建模方法

LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction

Jin Lou, Zhiyuan Jing, Xupeng Wang, Andong Chen, Xingdong Zhu, Yuexuan Li, Yuan Xu, Zhijie Zhu, Yingwei Ji, Wenpeng Nie, Renxing Feng, Liangliang Chen, Ying Chu, Jingyi Li, Jinyan Liu, Zhiqi Song, Jingxuan Zhu, Jidong Zhang, Yufei Liu, Boyang Xing, Lei Jiang, Yan Cui, Hongming Li, Yuchen Zhu

arXiv 2608.25757首次发表:更新:

发表机构

Humanoid Robot (Shanghai) Co., Ltd.; E-surfing Digital Life Technology Co., Ltd., China Telecom(人形机器人(上海)有限公司; 中国电信天翼数字生活科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LM-X是一种面向通用机器人操控的可解释动作建模方法,通过多尺度预测信号提升操控成功率,在RoboTwin2.0等任务上显著优于现有模型。

AI 中文摘要

通用视觉-语言-动作(VLA)策略主要通过短程动作预测学习长程行为,但仅能输出采样得到的指令,几乎无法提供额外信息。这造成两个耦合瓶颈:单一动作目标必须隐含地包含任务进度、中间意图与局部可靠性,而这些控制状态在执行过程中始终处于隐藏状态。受生物感知运动控制功能原理启发,我们提出LM-X,其在任务、事件与运动尺度上组织预测,不追求与生物结构的对应关系。模型在线生成三个显式监督信号,直接调控动作生成:待完成回报(RTG)衡量可见任务进度,待完成事件(ETG)识别下一个语义转换,异方差动作流通过传播方差估计局部可靠性。因此,解释是控制的固有属性,而非事后生成。在64块NVIDIA B200 GPU上进行耗时20天的预训练前,我们通过受控的五任务预训练门控验证了设计:完整模型的成功率仅动作骨干网络提升16.0个百分点,仅最强单头变体提升10.8个百分点。随后,我们在超过20000小时的真实机器人轨迹(含超过1000小时的失败策略 rollout)上训练LM-X。LM-X在50个随机困难RoboTwin2.0任务上的平均成功率达74.1%,而GR00T N1.7为55.4%;在7个真实机器人任务上为68.6%,GR00T N1.7为50.7%。RTG可追踪语义进度与可见退化,方差在犹豫与振荡控制期间上升。这些结果表明,显式多时间尺度预测状态可强化控制,同时提供可解释的内部估计。

英文摘要

Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: actions are exposed, but their explanatory state is not. They provide no native account of three explanatory signals: task progress, the next semantic transition, or local command reliability. Prior work shows that progress and event structure aid long-horizon control and that uncertainty supports monitoring; however, such capabilities are typically added or extracted only after action pretraining. The field therefore lacks a VLA foundation model whose explanatory state is jointly pretrained with control. Drawing on biological sensorimotor organization, in which outcome-sensitive, event-segmented, and probabilistic predictions structure behavior, we introduce LM-X. LM-X learns three directly supervised online signals: return-to-go (RTG) estimates visible progress and state quality; event-to-go (ETG) predicts the action sequence to the next semantic event; and heteroscedastic action-flow variance reports local command reliability. RTG conditions ETG and both condition action generation; uncertainty is estimated inside the action expert, making explanation part of control rather than a post-hoc description. We pretrain LM-X on more than 20,000 hours of heterogeneous real-robot trajectories, including over 1,000 hours of failed rollouts. LM-X achieves 74.1\% success on 50 randomized-hard RoboTwin2.0 tasks and 73.5\% on seven real-robot tasks, compared with 55.4\% and 50.7\% for GR00T N1.7. Its signals track progress and regression, anticipate event-scale motion, detect high-error actions, and provide advance failure warning. These results establish LM-X as an explainable VLA foundation model that couples transparent predictive state with stronger generalist control. Github: https://github.com/loongOpen/LoongWu-LM-X-VLA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑