arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越多模态对齐:通过响应替换与有序执行验证物理语言

Beyond Multimodal Alignment: Shared Physical Representations Across Sensors and Action Orders

Kaizhen Tan, Xin Xu, Siru Tao, Yixiao Li, Hanzhe Hong, Yang Feng, Heqing Du

arXiv 2608.19492首次发表:更新:

发表机构

New York University; Carnegie Mellon University; Columbia University(纽约大学; 卡内基梅隆大学; 哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出DBOSC方法,在Cluster Haptic数据集与弹塑性系统中验证了多模态物理表示的可执行含义,明确了执行器、图表等因素对动作组合的影响,区分了多项可测试的操作能力成就。

AI 中文摘要

世界模型日益将紧凑的多模态表示视为感知与物理交互的接口,但现有探测方法未明确不同传感器是否携带相同的可执行含义,或该含义能否在新的动作组合中保留。本文引入一种操作能力层级与不相交桥接算子替换证书(DBOSC),用于检验独立训练的模态编译器在训练集外的证据上是否可互换进入冻结响应图。在Cluster Haptic数据集上,相同未见过表面的音频与加速度表示在响应空间中的距离是错误表面配对的4.5倍,且该差距在所有19个未见过表面上均成立;对保留响应的解封证实,每个分支对物理的预测都优于总体图表。随后,本文在具有互补模态盲区的受控弹塑性系统中测试有序执行:在预注册预算下,先决条件拒绝该栈,因为冻结执行器无法通过未见过的程序推进甚至精确的图表坐标;在收敛预算下,相同的三阶图表可执行这些程序(神经常态均方误差NMSE为0.18),融合效果优于两种模态,16项注册检查中有14项通过;两项失败源于融合信息矩阵的对角限制与完整矩阵表现相当。通过该门限是执行器的属性,而非图表:在相同图表上,发射完整程序而非共享每步动态的执行器,比无实体预测器的表现差38倍。匹配的不可识别性结果解释了为何仅压缩与融合无法确定未见过的组合定律。这些结果将属性访问、响应替换、融合闭合与有序执行划分为可单独测试的不同成就。

英文摘要

Multimodal world models are often evaluated by whether different sensors produce similar representations. However, similar representations do not necessarily imply that the models make the same physical predictions, or that those representations can be reused when actions are combined in a new order. We study both questions through the physical responses predicted by a model. We first use the Cluster Haptic dataset to ask whether audio and acceleration can independently recover the behavior of the same surface from different observations. Predictions from the two sensors are substantially closer for the same surface than for different surfaces, with a $4.5\times$ gap on average, while both also outperform an average-surface prediction. We then show that this agreement alone does not determine how familiar actions should compose. In a controlled elastoplastic system, shared step dynamics fit observed programs less accurately than a whole-program predictor but generalize better to unseen action orders, with the ranking reversing on both held-out transitions across three independent initializations. Fusing free-decay and hysteresis observations further improves prediction, with diagonal Gaussian beliefs yielding the lowest errors. Together, these results distinguish cross-sensor consistency, multimodal fusion, and generalization to new action orders as separate questions in evaluating multimodal physical representations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑