超越视觉质量:世界动作模型测试时规划研究
Beyond Visual Quality: A Study of Test-Time Planning with World Action Models
浏览论文内容
中文总结 AI 辅助
本研究实证检验世界动作模型在测试时规划中的潜力,发现基于视觉质量等评分器仅能部分恢复选择机会,强调应通过决策有用性评估预测。
中文摘要 AI 辅助
世界动作模型生成动作及其后果的视觉预测。这些成对输出为规划创造了潜力:从一个状态采样多个动作,比较其想象结果,并选择预测结果最有前景的动作。然而,如何利用想象未来指导动作选择仍不清楚。我们实证检验了这一规划潜力。首先,我们通过选择实现结果最佳的采样候选来估计选择的上限。在受控的相同状态分析中,这一选择将成功率从均匀随机选择下的68.9%提升至79.2%。随后,我们测试了基于视觉质量、物理一致性和任务进展的选择器作为受控干预。一些测试的选择器产生了更高的观测成功率,但收益不均匀,且匹配的选择器未能恢复大部分已测量的机会。为调查这一差距,我们检验了采样动作是否导致不同结果、这些差异是否在预测中可见,以及评分是否能识别它们。从相同状态的反事实分支表明,选择机会集中在初始候选集中的少数决策中。动作扩散和结果覆盖不一定同步增加。在进一步评估中,跨越轨迹阶段并完整执行动作,测试的评分再次仅恢复了少量可用改进,尽管学习值带来小幅收益。这些发现区分了产生有后果的动作选择与在生成未来中识别它们,促使通过决策有用性而非仅视觉质量来评估WAM预测。
英文摘要
World action models generate actions together with visual predictions of their consequences. These paired outputs create the potential for planning by sampling multiple actions from one state, comparing their imagined outcomes, and choosing the action with the most promising predicted outcome. However, how to use imagined futures to guide action selection remains unclear. We examine this planning potential empirically. First, we estimate an oracle upper bound on selection by choosing the sampled candidate whose realised outcome is best. In a controlled same-state analysis, this choice raises success from 68.9% under uniform random selection to 79.2%. We then test selectors based on visual quality, physical consistency, and task progression as controlled interventions. Some tested selectors yield higher observed success, but the gains are uneven and the matched selectors leave much of the measured opportunity unrecovered. To investigate this gap, we examine whether sampled actions lead to different outcomes, whether these differences are visible in the predictions, and whether a score recognises them. Counterfactual branching from the same states shows that selection opportunity is concentrated in relatively few decisions in the initial candidate sets. Action spread and outcome coverage need not increase together. In a further evaluation across trajectory phases with complete action execution, the tested scores again recover little of the available improvement despite a small gain from learned value. These findings distinguish producing consequential action choices from recognising them in generated futures, motivating the evaluation of WAM predictions through their usefulness for decisions rather than visual quality alone.
发表机构
- University of Oxford(牛津大学)
- Purdue University(普渡大学)
- UWE Bristol(西英格兰大学布里斯托分校)
机构由 AI 辅助整理,请以论文原文为准。