arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解耦视觉-触觉前瞻:面向世界动作模型的神谕引导式接口发现

Disentangling Visuo-Tactile Foresight: Oracle-Guided Interface Discovery for World Action Models

Zihang Yao, Chaoyue Ding, Yingying Yu

arXiv 2608.00547首次发表:更新:

AI 中文总结

针对接触丰富的操纵任务中视觉与触觉跨模态接口未被充分探索的问题,提出OVTF框架与AFM模型,在UniVTAC基准7项任务中,AFM平均成功率优于IFM与UniVTAC-ACT。

AI 中文摘要

接触丰富的操纵任务仍具挑战性,因为成功的控制依赖于仅从视觉中往往难以观测到的物理交互线索。近期的触觉世界动作模型联合对未来视觉观测和触觉信号进行建模,以指导动作生成,但这类未来如何构建才能被动作专家有效利用的问题仍未得到充分探索。使用学习得到的世界动作模型直接研究该问题十分困难,因为端到端的行为会混杂物理上无效的视觉未来、不可靠的预测、不准确或跨模态不一致的触觉预测以及难以解读的未来到动作接口。为了使该接口可独立研究,我们引入了Oracle Visuo-Tactile Foresight(OVTF,神谕视觉-触觉前瞻),这是一个受控框架,它提供来自仿真中验证的成功轨迹的配对RGB和触觉未来。通过固定未来提供者,OVTF分离出接口并提出了一个更清晰的问题:如果未来是成功的且物理上可执行的,那么何种表示能让动作专家吸收其益处?在OVTF框架内,我们提出了Asymmetric Phase-Local Future Memory(AFM,非对称相位局部未来记忆),其中视觉记忆读取未来视觉,每个触觉记忆联合关注自身触觉流和相位对齐的未来视觉,且触觉间访问被阻断。我们将AFM与Modality-Isolated Future Memory(IFM,模态隔离未来记忆)进行比较,IFM移除了视觉到触觉的访问并独立处理每个未来模态。在UniVTAC仿真基准的7项任务中,AFM达到了32.0%的平均成功率,而IFM为23.7%,UniVTAC-ACT为14.9%。这项受控比较表明,选择性相位对齐的视觉-触觉路由相比完全模态隔离,能提供更具可操作性的未来到动作桥梁。

英文摘要

Contact-rich manipulation remains challenging because successful control depends on physical interaction cues that are often weakly observable from vision alone. Recent tactile world action models jointly model future visual observations and tactile signals to guide action generation, but how such futures should be structured for effective use by the action expert remains underexplored. Directly studying this question with learned world action models is difficult because end-to-end behavior entangles physically invalid visual futures, unreliable predictions, inaccurate or cross-modally inconsistent tactile forecasts, and an unreadable future-to-action interface. To make this interface independently studyable, we introduce Oracle Visuo-Tactile Foresight (OVTF), a controlled framework that supplies paired RGB and tactile futures from successful trajectories verified in simulation. By fixing the future provider, OVTF isolates the interface and asks a cleaner question: if the future is successful and physically executable, what representation allows the action expert to absorb its benefit? Within OVTF, we propose Asymmetric Phase-Local Future Memory (AFM), in which visual memory reads future vision, each tactile memory jointly attends to its own tactile stream and phase-aligned future vision, and cross-tactile access is blocked. We compare AFM with Modality-Isolated Future Memory (IFM), which removes visual-to-tactile access and processes each future modality independently. Across seven tasks on the UniVTAC simulation benchmark, AFM achieves 32.0% average success, compared with 23.7% for IFM and 14.9% for UniVTAC-ACT. This controlled comparison shows that selective phase-aligned visual-tactile routing provides a more actionable future-to-action bridge than complete modality isolation.

Comments6 pages, 3 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑