代数一致性本身并不能证明潜在动作模型中的时间结构
Algebraic Consistency Alone Does Not Certify Temporal Structure in Latent Action Models
浏览论文内容
中文总结 AI 辅助
本文证明潜在动作模型中的代数一致性(加法组合和反对称反转)并不能保证时间结构,因为重建驱动解码向状态特征差异,且约束模型在破坏配对后仍有效,建议采用基线校正和破坏配对重训练等验证协议。
中文摘要 AI 辅助
潜在动作模型推断无动作视频两帧之间转换的代码。近期方法正则化该代码以使其加法组合和反对称反转,并报告由此产生的误差数量级减少,作为代码捕获时间结构的无标签证明。我们表明这一结论并不成立。重建将解码后的转换推向状态特征的差异,对于任何配对,这两个恒等式都成立,该解是度量无法与编码干扰状态或坐标约定的解区分的。在五个源域中,一个训练过但无约束的对应模型已经实现了相对于未训练锚点的83-97%的减少。剩余倍数受解码器家族和所学内容的影响一样大。在其时间配对被破坏后重新训练的约束模型,在每个域中,在真实数据上仍达到比无约束模型更低的误差。在下游,保留时间配对在LIBERO-GOAL或LIBERO-SPATIAL上并未带来一致的优势,并且在测试的臂中,代码的平均线性动作可解码性随着代数误差的改善而下降。我们还测试了最直接的修复,即违反对比目标,要求代数在破坏的配对中失败:在测试配置中,它在重建预算内仅产生边际分离,在训练和测试三元组上都是如此。我们推荐这些方法目前缺乏的验证协议:基线校正度量、在破坏的配对上的重新训练以及种子预算分析。
英文摘要
Latent action models infer a code for the transition between two frames of action-free video. Recent methods regularise this code to compose additively and reverse antisymmetrically, and report order-of-magnitude reductions in the resulting errors as a label-free certificate that the code has captured temporal structure. We show that this conclusion does not follow. Reconstruction drives the decoded transition toward a difference of state features, for which both identities hold for any pairing, a solution the metric cannot distinguish from one encoding nuisance state or a coordinate convention. Across five source domains, a trained but unconstrained counterpart already achieves 83-97% of the reduction relative to an untrained anchor. The residual fold is governed as much by the decoder family as by what is learned. A constrained model retrained after its temporal pairing is destroyed still reaches, in each domain, a lower error than the unconstrained model on real data. Downstream, preserving the temporal pairing yields no consistent advantage on LIBERO-GOAL or LIBERO-SPATIAL, and across the tested arms the code's mean linear action decodability falls as the algebraic error improves. We also test the most direct repair, a violation-contrastive objective that requires the algebra to fail on destroyed pairings: in the tested configurations it yields only a marginal separation within the reconstruction budget, on training and test triples alike. We recommend a validation protocol that these methods currently lack: a baseline-corrected metric, retraining on destroyed pairings, and a seed-budget analysis.
发表机构
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
- Hunan University(湖南大学)
机构由 AI 辅助整理,请以论文原文为准。