arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LPA-CWM:一种用于反事实世界模型运动推理的学习型物理裁决器

LPA-CWM: A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models

Kunwei Wu, Xiang Liu, Guocai Yao, Junming Chen, Zhikang Chen, Min Zhang, Pengwei Wang, Sen Cui

arXiv 2609.14073首次发表:更新:

发表机构

National University of Singapore; Tsinghua University; Beijing Academy of Artificial Intelligence; Facebook; University of Oxford; East China Normal University(新加坡国立大学; 清华大学; 北京人工智能研究院; 脸书; 牛津大学; 华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LPA-CWM提出轻量级学习型物理裁决器,通过候选可靠性加权聚合反事实世界模型响应,在DAVIS和Kinetics上显著提升运动推理精度。

AI 中文摘要

反事实世界模型(CWM)通过比较事实预测与干预预测,从预训练的视频预测器中提取运动信息。然而,在不同目标帧掩码下生成的响应在可靠性上存在差异,而统一聚合则对其同等加权。我们将响应聚合问题表述为候选可靠性学习,并提出了LPA-CWM,其中包含一个轻量级的学习型物理裁决器(LPA)。该LPA在密集的MOVi-F轨迹上进行训练,拥有300万参数,通过比较无序候选集中的视觉上下文和响应结构来预测相对权重,而CWM预测器和干预生成器保持冻结。加权后的响应经过窗口化定位和一次配对重新评估以恢复运动。我们还引入了完整性感知的运动对应(CMC),这是一种以真实值为锚定的评估协议,联合衡量定位、轨迹完整性、可见性和连续性,将可见动态点上的缺失预测计为失败。在评估的DAVIS和Kinetics子集上,LPA-CWM相对于Uniform CWM将DCA_avg分别提高了60.0%和29.0%,并且在TAP-Vid First基准下也提高了跟踪精度。快速概览可访问此https URL。

英文摘要

Counterfactual world models (CWM) extract motion from pretrained video predictors by comparing factual and intervened predictions, but uniform aggregation weights responses equally without explicitly incorporating physical priors. Our key insight is to incorporate physical priors into candidate reliability learning, motivating LPA-CWM with a lightweight Learned Physical Adjudicator (LPA). Trained on dense MOVi-F trajectories, the 3.0M-parameter LPA compares visual context and response structure across an unordered candidate set to predict relative weights; windowed localization and one paired re-evaluation recover motion with the CWM frozen. Existing video-level benchmarks do not directly assess motion correspondence, where low localization error can conceal missing trajectory segments. We introduce Completeness-aware Motion Correspondence (CMC), a ground-truth-anchored protocol jointly measuring localization, completeness, visibility, and continuity, counting missing predictions as failures on visible dynamic points. Across DAVIS, Kinetics, and RoboTAP, LPA-CWM improves all main CMC measures over Uniform CWM, with relative gains of 18.1%--60.0% in average Dynamic Correspondence Accuracy ($\mathrm{DCA}_{\mathrm{avg}}$), and improves TAP-Vid First tracking accuracy (overview: https://LPA-CWM.github.io).

CommentsA quick overview is available at https://LPA-CWM.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑