arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估带点质量数据上的生成式时间序列模型

Evaluating Generative Time-Series Models on Data with Point Masses

Jian Xu

arXiv 2608.09692首次发表:更新:

发表机构

RIKEN iTHEMS; RIKEN Center for Advanced Intelligence Project (AIP)(理化学研究所 iTHEMS; 理化学研究所先进智能项目中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对带点质量的时间序列数据,发现现有基准协议存在原子结构不匹配等问题,通过控制方法与多模型基准测试,揭示了时间耦合的作用及不同模型在发生统计量下的排序差异。

AI 中文摘要

许多用于生成式时间序列模型基准测试的序列,会将大量概率质量集中在单个值上——比如“不下雨”“无订单请求”“未订购零件”这类情况。本文报告了对这类数据进行仔细评估时的发现:首先,标准的滚动原点协议可能在原子结构与数据集毫无相似之处的窗口上对模型评分,在一个基准数据集中,42%的值为零,而评估窗口中仅13%;另一个基准数据集47%的值为零,评估窗口仅5%。这并非无关紧要的问题——它推翻了我们自己的一项结论,将研究中表现最强的发生模型变成了警示案例。其次,本文提供了一种控制方法,其中CRPS(连续排名概率得分)在构造上具有不变性,同时时间耦合被破坏,这能精确测量该耦合对所选统计量的贡献程度。第三,在5个随机种子下,采用匹配协议对7个模型进行基准测试,自回归 hurdle 模型在6个数据集中的5个上优于条件流模型,差距最高达153倍;而流模型自身的发生统计量在不同训练种子间的变化高达62%,且所有基线模型均为确定性模型。最后,在5种不同的发生统计量下,模型排序并不相同,且构造不同的两个模型间一致性最低。

英文摘要

Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not rain, no ride is requested, no part is ordered. We report what happens when such data is evaluated carefully. First, the standard rolling-origin protocol can score a model on a window whose atom structure bears no resemblance to the dataset: on one benchmark the dataset is $42\%$ zeros and the evaluation windows are $13\%$, on another $47\%$ against $5\%$. This is not a cosmetic problem --- it reversed one of our own conclusions, turning the strongest occurrence model in our study into what looked like a cautionary tale. Second, we give a control in which CRPS is invariant \emph{by construction} while the temporal coupling is destroyed, which measures exactly how much that coupling contributes to a chosen statistic. Third, benchmarking seven models on a matched protocol over five seeds, an autoregressive hurdle beats a conditional flow on five of six datasets, by up to a factor of $153$, while the flow's own occurrence statistics vary by up to $62\%$ across training seeds and every baseline is deterministic. Finally, the model ordering is not the same under five different occurrence statistics, and the two that do not share a construction agree with each other least.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑