AI 中文总结
该研究探讨时间序列世界模型中预测准确性与机制一致性的分歧,提出形式化定义和基准测试,发现冻结潜在空间和输出融合提升准确性,但机制一致性需方向监督训练目标。
AI 中文摘要
时间序列世界模型(TSWM)根据观测到的历史、计划动作和外部输入来预测受控系统的状态。当前方法将动作作为协变量构建预测器,并在执行计划下的预测误差上进行训练和评估。然而,世界模型比较的是未执行计划,但它们对变化计划的响应仍未得到测试。我们探究哪些设计选择至关重要,以及准确的预测器是否像真实系统一样对变化计划作出响应。我们通过形式化和基准测试来解决这两个问题。形式化将状态、动作和外部输入分离,区分连续、模式和事件动作,并引入机制一致性,这是一种基于已声明动作-状态关系(具有已知方向,如血管升压药升高血压)构建的度量:它检查移动动作是否使预测沿声明方向移动。基准测试整合了八个公共数据集,包含来自工程基础设施和临床护理的真实动作,在七个骨干网络和五个随机种子上变化预测空间、计划融合和计划编码。首先,冻结的潜在预测空间相比观测空间平均降低MAE 9.9%,门控输出融合相比输入拼接平均降低12.7%,两者均改善了全部八个数据集;时间计划编码平均MAE变化最多2.2%。其次,预测误差和机制一致性出现分歧:在五个具有声明机制的数据集中,最低误差配置在四个数据集上的一致性达到或低于随机水平,且没有设计选择能避免这一点。最后,方向监督,一种对移动动作响应中错误符号部分进行惩罚的损失,显著提高了受惩罚机制的一致性,且MAE不变。这些共同为TSWM提供了一个配方:冻结的潜在空间和输出侧融合用于准确性,以及一个用于机制一致性的训练目标。
英文摘要
A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but their responses to changed plans remain untested. We ask which design choices matter and whether accurate forecasters respond to changed plans as real systems do. We address both with a formalization and benchmark. The formalization separates state, actions and exogenous inputs, distinguishes continuous, mode and event actions, and introduces mechanism consistency, a metric built on declared action-state relations with known directions, such as a vasopressor raising blood pressure: it checks whether shifting an action moves the forecast in the declared direction. The benchmark consolidates eight public datasets with real actions from engineered infrastructure and clinical care, varying prediction space, plan fusion and plan encoding across seven backbones and five seeds. First, a frozen latent prediction space lowers MAE by 9.9% over observation space and gated output fusion lowers it by 12.7% over input concatenation on average, with both improving all eight datasets; temporal plan encoding changes average MAE by at most 2.2%. Second, prediction error and mechanism consistency diverge: the lowest-error configuration is at or below chance in consistency on four of five datasets with declared mechanisms, and no design choice avoids this. Finally, directional supervision, a loss penalizing the wrong-signed part of the response to a shifted action, significantly raises consistency on penalized mechanisms with no change in MAE. Together they give TSWMs a recipe: a frozen latent space and output-side fusion for accuracy, and a training objective for mechanism consistency.