arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30721cs.LG

当10000个窗口并非10000次测试:审计滑动窗口时间序列分类中的统计置信度

When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification

  • Beijing University of Posts and Telecommunications(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

Xinze Shi, Litian Zhang, Binrui Shi

AI总结:

针对滑动窗口时间序列分类中重叠测试窗口非独立的问题,提出审计方法映射三种声明至依赖稳健推断,揭示信息增长远低于窗口增长,并给出可检查的工作流程。

AI中文摘要:

滑动窗口分类器通常在数千个重叠的测试窗口上进行评估,尽管相邻预测共享观测值,并且仍然嵌套在记录和受试者内部。受试者不相交的评估防止了一种形式的泄漏,但并未使这些测试窗口独立。我们提出了一种实用的审计方法,将三个声明——对已观测记录的性能、对已观测受试者的未来记录的性能以及对未见受试者的性能——映射到明确的聚合规则和成熟的依赖稳健推断。在75%重叠下,受控模拟对IID已观测记录推断给出16.9%的I类错误,对会话中心化Bartlett-HAC给出7.2%:这是一个显著的改进,但仍有残余的校准偏差。对冻结的WISDM和HARTH预测的审计表明,测试行数增长近四倍仅带来1.75-1.94倍的方差等效信息增长。在该重叠水平下,固定记录的配对准确率差异区间宽度是IID区间宽度的1.22-1.66倍;这种膨胀在零重叠时并非普遍存在。在HARTH上,配对准确率差异区间在三种重叠设置下均包含零,而Macro-F1则偏向MiniROCKET。独立重新计算、共同会话检查、类别级结果以及分别播种的校准使得审计的范围和局限性可被检查。由此产生的工作流程区分了额外预测与额外独立证据。

英文摘要:

Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does not make those test windows independent. We present a practical audit that maps three claims - performance on observed recordings, future recordings from observed subjects, and unseen subjects - to explicit aggregation rules and established dependence-robust inference. At 75% overlap, controlled simulations give 16.9% Type-I error for IID observed-record inference and 7.2% for session-centered Bartlett-HAC: a substantial improvement with residual miscalibration. Audits of frozen WISDM and HARTH predictions show that nearly fourfold growth in test rows yields only 1.75-1.94-fold variance-equivalent information growth. At that overlap, fixed-record paired Accuracy-difference intervals are 1.22-1.66 times the IID widths; this inflation is not universal at zero overlap. On HARTH, paired Accuracy-difference intervals include zero across three overlap settings, whereas Macro-F1 favors MiniROCKET. Independent recomputation, common-session checks, class-level results, and separately seeded calibration make the audit's scope and limitations inspectable. The resulting workflow distinguishes additional predictions from additional independent evidence.

补充信息

↑