arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DoTime:用于干预和反事实时间序列的合成基准生成器

DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

Dennis Thumm, Billy Tim Anthony, Ying Chen

arXiv 2607.27263首次发表:更新:

发表机构

National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DoTime是用于干预和反事实时间序列的合成基准生成器,具备现有工具没有的多项功能,附带评估套件与基线,经测试其干预训练的因果模型相比同容量观测模型有方向准确性优势。

AI 中文摘要

大多数用于时间序列因果推断的基准是观测性的、小规模的或特定领域的,而在医疗保健、政策评估和气候科学等最重要的场景中,干预和反事实估计却未得到充分支持。我们推出DoTime,这是一个开源、可扩展且具有理论基础的多变量时间结构因果模型(TSCM)生成器,以dotime PyPI包的形式发布,附带四个固定的评估套件。与现有工作相比,它具备现有生成器所没有的功能:连续时间干预窗口、带有正性保护的反事实采样模式、作为中断时间序列严格泛化的机制切换SCM、通过切换SCM参数构建的非平稳动态,以及将趋势和结构断点置于评估窗口内的确定性斜坡和正弦干预轮廓。此外,它证明了该生成器可作为因果基础模型参考实现的先验。发布的套件包含100000条轨迹的训练规模快照和八种命名识别结构,每种结构都有精确的ground truth:整个过程来自同一SCM的配对干预轨迹,以及连续时间套件中的共享噪声反事实。我们附带了带有评估工具的参考基线实现,并提出了一个可证伪的主张:干预训练相较于相同容量的观测模型,具有可测量的方向准确性优势。该主张在每组三个训练种子上进行了测试,在针对保留的episode进行结构匹配评估时,干预先验拟合网络(PFN)的差距在所有测试的结构、轨迹长度和种子中均为正。

英文摘要

Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy evaluation, and climate science. We introduce \textbf{DoTime}, an open, scalable, and theoretically grounded generator of multivariate temporal structural causal models (TSCMs) with interventions, released as the \code{dotime} PyPI package together with four frozen evaluation suites. Beyond existing work, it adds capabilities absent from prior generators: continuous-time intervention \emph{windows}, counterfactual sampling modes with a positivity guard, regime-switching SCMs as a strict generalization of interrupted time series, non-stationary dynamics by construction with switching SCM parameters, and deterministic ramp and sinusoidal intervention profiles that place trends and structural breaks \emph{inside} the evaluation window. Moreover, it demonstrates the suitability of the generator as a prior for a causal foundation model reference implementation. The released suites span a training-scale snapshot of $100{,}000$ trajectories and eight named identification structures, each with exact ground truth: paired interventional trajectories from the same SCM throughout, and shared-noise counterfactuals in the continuous-time suite. We ship reference baseline implementations with an evaluation harness, and pose a falsifiable claim: interventional training buys a measurable direction-accuracy advantage over an observational model of identical capacity. It is tested across three training seeds per arm. Under structure-matched evaluation on held-out episodes, the interventional prior-fitted network's (PFN) gap is positive in every structure, trajectory length, and seed tested.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑