发表机构
State Key Laboratory of CAD&CG, Zhejiang University; Institute for AI Industry Research (AIR), Tsinghua University; InSpatio(浙江大学计算机辅助设计与图形学国家重点实验室; 清华大学人工智能产业研究院; InSpatio)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对前向潜在世界模型恢复动作需昂贵搜索的问题,提出INTACT方法,通过特定架构和技术实现意图到动作学习,在多任务上取得高成功率,减少采样并提升性能,还具有快速推理能力。
AI 中文摘要
前向潜在世界模型预测动作如何改变场景,但仅通过昂贵的测试时间搜索来恢复期望变化的动作。我们引入了INTACT(意图到动作),这是一种端到端的JEPA,它将有动作标签、无奖励的轨迹转化为可部署的意图到动作接口。每个转换提供物理意图$z_{t + 1}-z_t$,而未来目标提供部署意图$\text{sg}(z_g)-z_t$。该架构通过相同的四槽语法和共享参数在局部和目标运动意图主干输入图之间是同构的,并且通过相同预测器诱导的动作定律语义在支持的局部和目标运动意图族之间是同构的。INTACT还提供从RGB证据到动作有效潜在意图坐标以及从意图族到其相应动作定律族的完整转移。不对称端点梯度为物理后继提供基础并将未来目标固定为锚点,在没有逐点潜在匹配或全局线性动力学的情况下结合表示学习和控制。所得坐标支持强大的分布动作定律:其条件均值直接用作无搜索策略,同时采样可用于多样性或可选验证。在四个官方LeWM任务上,单轮、零搜索模型的成功率达到85.78%、100.00%、97.67%和97.89%。以直接计划为中心的可选局部CEM使用384个而不是9000个候选序列达到96.86%的宏观成功率,采样减少了23.44倍,同时将纯CEM提高了16.00分。一个共享的四任务编码器达到89.39%的E5直接宏观成功率,并在联合训练的LeWM上改进了每个任务,而预测专家动作族kNN在$r = 0.954$时跟踪直接成功率。直接推理需要2.9 - 5.5毫秒。
英文摘要
Forward latent world models predict how actions change a scene, but recover actions for a desired change only through expensive test-time search. We introduce INTACT (INtent-To-ACTion), an end-to-end JEPA that turns action-labeled, reward-free trajectories into a deployable intent-to-action interface. Each transition supplies physical intent $z_{t+1}-z_t$, while a future goal supplies deployment intent $\operatorname{sg}(z_g)-z_t$. The architecture is isomorphic between the local and goal motion-intent backbone-input graphs through an identical four-slot grammar and shared parameters, and between supported local and goal motion-intent families through action-law semantics induced by the same predictor rather than pointwise latent equality. INTACT also provides intact transfer from RGB evidence to action-effective latent intent coordinates and from intent families to their corresponding action-law families. Asymmetric endpoint gradients ground physical successors and fix future goals as anchors, joining representation learning and control without pointwise latent matching or globally linear dynamics. The resulting coordinates support a robust distributional action law: its conditional mean serves directly as a search-free policy, while sampling remains available for diversity or optional verification. On the four official LeWM tasks, one-epoch, zero-search models reach 85.78\%, 100.00\%, 97.67\%, and 97.89\% success. Optional local CEM centered on the Direct plan reaches 96.86\% macro success using 384 instead of 9,000 candidate sequences, reducing sampling by $23.44\times$ while improving pure CEM by 16.00 points. One shared four-task encoder reaches 89.39\% E5 Direct macro and improves every task over jointly trained LeWM, while predicted--expert action-family kNN tracks Direct success at $r=0.954$. Direct inference takes 2.9--5.5 ms.
Comments28 pages, 11 figures, including appendices