arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11875cs.RO

UniMPA:一种通过动作接地转移建模的统一记忆-预测-动作模型

UniMPA: A Unified Memory-Prediction-Action Model via Action-Grounded Transition Modeling

发表机构哈尔滨工业大学(深圳) · 南洋理工大学
查看机构详情
  • Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Wei Li, Rui Shao, Jie He, Lingsen Zhang, Ziwei Liu, Liqiang Nie

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA模型中的转移可实现性差距问题,提出UniMPA统一模型,通过动作接地转移接口、持久选择性预测和双记忆库检索,解决转移模糊性、预测执行不匹配及经验实现不匹配。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型的最新进展改善了机器人操作,然而从观察到动作的学习仍然受到一个基本的转移可实现性差距的限制,这一差距体现在三个紧密耦合的问题上:(i)转移模糊性。视觉上相似的当前观察可能对应于不同的操作阶段,并暗示不同的后续转移。(ii)预测与执行不匹配。视觉上合理的预测未来观察不一定对应于物理上可实现的转移。(iii)经验与实现不匹配。历史上可执行的动作模式不一定能在当前场景中实现预期的转移,因此需要上下文感知的适应。据此,我们提出了UniMPA,一种统一的记忆-预测-动作模型,通过共享的动作接地转移接口来解决这些问题。(i)UniMPA引入了持久选择性未来预测,通过建模预期的未来状态演化来解决转移模糊性。一个持久潜在流持续跟踪任务级进展,而一个转移关键像素流通过记忆接地预测选择性地解决细粒度交互变化。(ii)为了评估预期转移的物理可执行性,预测的转移查询一个时间视觉-动作记忆库。该库检索历史上实现的视觉-动作经验,将未来预测接地于可执行的证据。(iii)为了将可执行经验适应到当前场景,一个动作-视觉记忆库从历史动作演化中检索视觉接地的动作原型。原型偏置流然后将流源移向历史支持的动作流形,以进行上下文感知的细化。

英文摘要

Recent advances in Vision-Language-Action (VLA) models have improved robotic manipulation, yet observation-to-action learning remains limited by a fundamental transition realizability gap, manifested in three tightly coupled problems: (i) Transition ambiguity. Visually similar current observations may correspond to different manipulation phases and imply different subsequent transitions. (ii) Prediction--execution mismatch. A visually plausible predicted future observation does not necessarily correspond to a physically realizable transition. (iii) Experience--realization mismatch. A historically executable action pattern may not necessarily realize the intended transition in the current scene and therefore requires context-aware adaptation. Accordingly, we propose UniMPA, a Unified Memory-Prediction-Action model that addresses these problems through a shared action-grounded transition interface. (i) UniMPA introduces Persistent-Selective Future Prediction to resolve transition ambiguity by modeling the intended future state evolution. A persistent latent stream continuously tracks task-level progress, while a transition-critical pixel stream selectively resolves fine-grained interaction changes through memory-grounded prediction. (ii) To assess the physical executability of the anticipated transition, the predicted transition queries a temporal Visual-Action Memory Bank. The bank retrieves historically realized visual-action experience, grounding future prediction in executable evidence. (iii) To adapt executable experience to the current scene, an Action-Visual Memory Bank retrieves visually grounded action prototypes from historical action evolution. Prototype-Biased Flow then shifts the flow source toward a historically supported action manifold for context-aware refinement.

补充信息

↑