arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越输出空间校准:用于时间序列分类中选择性可靠性估计的频谱证据捆绑

Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification

Filippo Cenacchi, Longbing Cao, Runze Yang

arXiv 2607.18279首次发表:更新:

发表机构

Macquarie University(麦考瑞大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对时间序列分类可靠性问题,提出验证门控固定标签可靠性策略,结合输出端线索与频谱描述符形成可靠性估计,经实验在多个数据集和骨干家族上提升了选择性可靠性指标,验证门控进一步优化了结果。

AI 中文摘要

时间序列分类的事后校准通常重新映射输出分数,但诸如信任、弃权和审查等部署决策取决于当前时间信号是否支持自信的预测。我们解决了三个时间序列可靠性差距:相同的置信值可能隐藏不同的时间支持,平均校准可能错过错误的高置信度误差,输出空间重新校准提供的输入关联可审计性有限。我们引入了一种验证门控固定标签可靠性策略,在估计是否应信任骨干预测时保持其不变。该方法将输出端线索与全样本频谱描述符相结合,形成标量可靠性估计和诊断频段级证据。验证门仅在正确性排名提高且不违反FalseConf@0.9或AURC容限时启用频谱调节;否则恢复到更安全的输出空间基线。在八个异构UCR/UEA数据集、八个时间序列骨干家族和标准重新校准器上,无约束方法提高了匹配评估子集上的固定标签选择性可靠性指标,将Corr - AURC从0.693提高到0.779。验证门控策略进一步将Corr - AURC提高到0.786,并将FalseConf@0.9降低到0.094。这些结果表明,时间序列分类器的可靠性估计受益于将输出置信度与频谱证据捆绑,而验证门控可防止无支持的频谱调节。

英文摘要

Post-hoc calibration for time-series classification usually remaps output scores, but deployment decisions such as trust, abstention, and review depend on whether a confident prediction is supported by the current temporal signal. We address three time-series reliability gaps: identical confidence values can hide different temporal support, average calibration can miss false high-confidence errors, and output-space recalibration offers limited input-linked auditability. We introduce a validation-gated fixed-label reliability policy that keeps the backbone prediction unchanged while estimating whether it should be trusted. The method combines output-side cues with whole-sample spectral descriptors, including band energy, entropy, peak dominance, period support, and phase stability, to form a scalar reliability estimate and diagnostic band-level evidence. A validation gate enables spectral conditioning only when correctness ranking improves without breaching FalseConf@0.9 or AURC tolerances; otherwise it reverts to the safer output-space baseline. Across eight heterogeneous UCR/UEA datasets, eight time-series backbone families, and standard recalibrators, the unconstrained method improves fixed-label selective-reliability metrics on the matched evaluation subset, raising Corr-AURC from 0.693 to 0.779. The validation-gated policy further improves Corr-AURC to 0.786 and reduces FalseConf@0.9 to 0.094. These results suggest that reliability estimation for time-series classifiers benefits from bundling output confidence with spectral evidence, while validation gating prevents unsupported spectral conditioning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑