arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有误差都重要:决策相关预测误差预测规划质量

Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality

Linhao Wang, Yiyan Fan, Dongjin Huang

arXiv 2609.32322首次发表:更新:

AI 中文总结

本研究提出决策相关预测误差(DRPE),证明其比总预测误差更能预测规划质量,并通过等误差评估协议在网格世界中验证,DRPE与规划成功率强相关(ρ=-0.84),而总误差相关性弱(ρ=-0.25)。

AI 中文摘要

世界模型通常通过预测误差进行训练和评估,假设更准确的预测会带来更好的决策。我们表明这一假设可能失效,因为总误差相似的模型,当其误差发生在不同的状态维度上时,规划性能可能差异显著。我们引入决策相关预测误差(DRPE),该指标衡量影响决策的状态维度上的预测误差。我们还开发了一种等误差评估协议,在保持总误差固定的同时改变误差分配。在一个具有已知状态相关性和标准化规划器的因子化网格世界中,我们评估了55个受控和学习的模型,涵盖不同的误差水平和分配。总预测误差与规划成功率的相关性较弱(Spearman ρ=-0.25),而DRPE则具有很强的预测性(ρ=-0.84;在受控模型族内为-0.98)。总误差仅相差1%的模型,其规划成功率可相差60个百分点(97%对37%)。相关误差还取决于任务,在相同总误差下,模型排名在不同任务间发生反转。更深的想象进一步放大决策相关误差,而学习模型在罕见但决策关键的事件上表现出系统性偏差。我们形式化了DRPE能正确排序模型而总预测误差不能的充分条件。

英文摘要

World models are typically trained and evaluated by prediction error, assuming that more accurate predictions lead to better decisions. We show that this assumption can fail because models with similar total error can differ substantially in planning performance when their errors occur on different state dimensions. We introduce Decision-Relevant Prediction Error (DRPE), which measures prediction error on the state dimensions that affect decisions. We also develop an iso-error evaluation protocol that varies error allocation while keeping total error fixed. In a factored gridworld with known state relevance and a standardized planner, we evaluate 55 controlled and learned models across different error levels and allocations. Total prediction error is weakly related to planning success (Spearman $ρ=-0.25$), whereas DRPE is strongly predictive ($ρ=-0.84$; $-0.98$ within the controlled family). Models with only a 1\% difference in total error can differ by 60 percentage points in planning success (97\% vs 37\%). The relevant error also depends on the task, with model rankings reversing across tasks at the same total error. Deeper imagination further amplifies decision-relevant errors, while learned models exhibit systematic bias on rare but decision-critical events. We formalize sufficient conditions under which DRPE correctly ranks models and total prediction error cannot.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑