当鞅永不停止触发:真实预测流上的任意时刻有效门控机制
When the Martingale Never Stops Firing: Anytime-Valid Gating on Real Forecast Streams
浏览论文内容
中文总结 AI 辅助
本研究针对真实预测流场景,测试了任意时刻有效门控机制在修正时间序列基础模型时的有效性,发现其在真实流上触发异常,采用Huber式门控可显著降低退化,建议相关方法附带空校准控制。
中文摘要 AI 辅助
机器学习系统在运行过程中不断被修正,何时进行干预的决策越来越多地委托给统计监控器。任意时刻有效推理提供了可在任何时刻采取行动的证据,这正是该场景所需的保证,且该方法正从理论走向部署后的监控。共形检验鞅是变化检测工具,Ville不等式限制了其在可交换数据上的虚警概率,该保证是有条件的:仅当被监控的流表现为可交换时,部署的监控器才继承该保证。这一前提在监控器最有用的场景(即依赖数据及监控器修改其读取分数的学习器的循环内部)中最难满足,且很少被测量。我们在预先指定的案例研究中对其进行测量:某监控器对卡尔曼适配器的在线更新进行门控,该适配器用于修正五个预测流上的冻结时间序列基础模型。在可交换合成流上,相同实现的触发次数最多为60次运行中的1次;在真实流上,当α=0.05时,135次干净流运行全部触发。该构造无法解释触发原因,失败源于部署的分数流本身:重复触发使门控的漂移响应保持活跃,而门控滤波器放大了其本应预防的瞬态,值得保留的组件不做有效性声明。对滤波器自身更新采用Huber式门控,可使孤立尖峰的退化降低一个数量级,且无需特定数据集调整。因此,针对依赖数据提出的任意时刻有效方法应附带空校准控制和机制轨迹。
英文摘要
Machine learning systems are increasingly corrected while they run, and the decision of when to intervene is increasingly delegated to statistical monitors. Anytime-valid inference promises evidence that can be acted on at any moment, exactly the guarantee this setting needs, and it is moving from theory into deployed monitoring. Conformal test martingales are the change-detection instrument, and Ville's inequality caps their false-alarm probability on exchangeable data. The guarantee is conditional. A deployment inherits it only if the stream it monitors behaves exchangeably. The premise is hardest to satisfy where these monitors are most useful, on dependent data and inside loops where the monitor modifies the learner whose scores it reads. It is also rarely measured. We measure it in a pre-specified case study, where such a monitor gates the online updates of a Kalman adapter correcting frozen time-series foundation models on five forecasting streams. On exchangeable synthetic streams, the same implementation fires in at most 1 of 60 runs. On the real streams, at alpha = 0.05, 135 of 135 clean-stream runs fired. The construction does not explain the firing; the failure comes from the deployed score stream itself. Repeated fires hold the gate's drift response active, and the gated filter amplifies the very transient it was designed to prevent. The component worth keeping makes no validity claim. Huber-style gating of the filter's own updates cuts isolated-spike degradation by an order of magnitude with no dataset specific tuning. Anytime-valid methods proposed for dependent data should therefore be accompanied by null-calibration controls and mechanism traces.
发表机构
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。