arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TFP:用于视觉运动学习的时间条件记忆融合策略

TFP: Temporally Conditioned Memory-Fusion Policies for Visuomotor Learning

Yushen Liang, Yue Peng, Baosheng Jin, Tianluo Zhang, Xinyu Zhang, Shuyi Zhou, Zhuoran Chen, Xinqi Liu, Shenji Wan

arXiv 2607.08283首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Beijing Institute of Technology(电子科技大学; 北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视觉运动学习中阶段依赖操作问题,提出时间条件记忆融合策略(TFP),通过维护任务进展信念并注入动作解码器来改进VLA策略,在多个任务上提升成功率,证明紧凑且对事件敏感的记忆动态可优化VLA策略。

AI 中文摘要

视觉-语言-动作(VLA)策略在许多操作任务中表现良好,但往往具有反应性,根据当前观察、指令和本体感觉状态预测下一个动作。在阶段依赖的操作中,这种假设不成立。我们认为此类任务不仅需要记忆,还需要动态感知的信念更新。我们引入了时间条件记忆融合策略(TFP),这是一个用于VLA主干的轻量级记忆-动作框架。TFP通过液体时间常数动态维护局部情节任务进展信念,并通过自适应调制将更新后的信念直接注入流匹配动作解码器。在LIBERO和LIBERO-plus上,TFP提高了平均成功率。在以记忆为重点的MIKASA ShellGameTouch诊断中,TFP成功率高达75.0%。机械分析表明,操作事件附近的写入增益变化比非事件阶段大6倍左右,隐藏状态干预表明信念因果调节生成的动作块。这些结果表明,紧凑、对事件敏感的记忆动态可以改善VLA策略。

英文摘要

Vision--Language--Action (VLA) policies such as $π_{0.5}$ and OpenVLA perform well on many manipulation tasks, but they are often reactive: the next action is predicted from the current observation, instruction, and proprioceptive state. This assumption breaks down in stage-dependent manipulation, where visually similar states may require different actions depending on latent task progress and previous interaction outcomes. We argue that such tasks require not only memory, but dynamics-aware belief updates: the policy should preserve task progress during stable or occluded phases and revise its belief near contact, release, or subgoal transitions. We introduce Temporally Conditioned Memory-Fusion Policies (TFP), a lightweight memory-action framework for VLA backbones. TFP maintains an episode-local task-progress belief with Liquid Time-Constant dynamics and injects the updated belief directly into the flow-matching action decoder through adaptive modulation. This lets temporally accumulated context shape the generated action chunk, rather than serving only as passive history context. With a 3.3B-parameter model, TFP improves the average success rate from 96.9% to 98.75% on LIBERO and from 91.4% to 93.77% on LIBERO-plus. On the memory-focused MIKASA ShellGameTouch diagnostic, TFP achieves success up to 75.0%. Mechanistic analyses show that write-gain changes near manipulation events are about 6 times larger than far non-event phases, and hidden-state interventions show that the belief causally modulates generated action chunks. These results suggest that compact, event-sensitive memory dynamics can improve VLA policies under occlusion, visual perturbation, and stage-dependent task structure.

CommentsAccepted to the SemRob 2026 Workshop at Robotics: Science and Systems (RSS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑