发表机构
Huawei; University of Science and Technology of China; Zhejiang University; Tsinghua University; Harbin Institute of Technology; Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(华为; 中国科学技术大学; 浙江大学; 清华大学; 哈尔滨工业大学; 广东省人工智能与数字经济实验室(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对JEPAs世界模型易受视觉扰动影响的问题,提出动作条件预测一致性(ACPC)诊断方法,定义IR与SR指标,经四视觉控制任务实验验证其可预测扰动带来的预测及代价变化。
AI 中文摘要
联合嵌入预测架构(JEPAs)学习世界模型时在紧凑的潜在空间而非像素空间中进行预测,降低了对干扰外观建模的压力,但这无法保证抵御视觉扰动:视觉扰动仍可改变编码表示并影响后续动作条件预测。双模拟(Bisimulation)精确捕捉了这一要求:仅当两个观测值的动作条件结果一致时,才应将它们视为同一状态。基于该准则,我们提出动作条件预测一致性(ACPC),这是一种诊断方法,用于衡量干净历史及其视觉扰动视图在相同动作序列下向前推演后出现的偏差程度。我们证明该偏差可界定扰动诱导的多步预测误差和规划器代价的变化。基于成对ACPC,我们定义了两个互补指标:不变半径(IR)总结干净-扰动推演的偏差范围,分离率(SR)检查不同状态在推演后是否仍可区分。在四个视觉控制任务上的实验表明,成对ACPC可预测扰动诱导的预测和代价变化;在LeWM数据集上,IR-SR筛选可跨任务迁移,且该联合诊断在模糊和缩放场景下仍有效;PLDM在不同架构下表现出相似的诊断趋势。
英文摘要
Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance. Yet this provides no guarantee against visual perturbations: they can still alter the encoded representation and affect subsequent action-conditioned predictions. Bisimulation captures this requirement precisely: two observations should be treated as the same state only when their action-conditioned consequences agree. Guided by this criterion, we introduce Action-Conditioned Predictive Consistency (ACPC), a diagnostic that measures how far a clean history and a visually perturbed view of it diverge after being rolled forward under the same action sequence. We prove that this divergence bounds the perturbation-induced change in multi-step prediction error and planner cost. Building on pairwise ACPC, we define two complementary measures: the Invariance Radius (IR) summarizes clean-perturbed rollout spread, while the Separation Rate (SR) checks whether different states remain distinguishable after rollout. Experiments on four visual control tasks show that pairwise ACPC predicts perturbation-induced prediction and cost changes. On LeWM, the IR-SR screen transfers across tasks, and the joint diagnostic remains informative under blur and resize. PLDM exhibits similar diagnostic trends under a different architecture.