arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08424stat.MLcs.CVcs.LG

ARC:用于变点定位的增广秩共形化方法——有限样本有效性与分布鲁棒效率

ARC: Augmented-Rank Conformalization for Changepoint Localization --- Finite-Sample Validity and Distribution-Robust Efficiency

Chenchen Peng, Mixia Wu, Qijing Yan, Zhiqi Shen, Jie Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出ARC增广秩共形化方法,可实现变点定位的有限样本覆盖率,且在单调变换下置信集长度稳定,经仿真与well-log基准测试验证了其有效性。

中文摘要 AI 辅助

共形变点定位可将任意得分转化为具有有限样本覆盖率的变点置信集,其覆盖率具有普适性,但效率并非如此。最优得分是似然比,因此实际应用中的得分需估计密度比,且在重尾、偏态及分布偏移场景下,置信集长度会恶化,此时无长度保证。本文提出ARC(增广秩共形化),这是一类仅通过段内秩依赖数据的得分,包括秩-CUSUM位置与尺度通道、它们的固定组合,以及经合成训练后冻结的轻量神经得分。每个ARC得分在所有冻结权重配置(包括随机初始化与训练失误)下均继承有限样本覆盖率。核心结果为效率传递定理:整个ARC置信集在严格单调边际变换下几乎必然不变,因此置信集长度分布仅通过秩结构依赖于数据对,且某一时刻的长度保证在其单调轨道上逐字成立,而插件得分的长度会随每次重表达而变化。不同秩结构下长度确实会改变,并如实报告。经典秩检验理论将ARC定位为以有界成本实现最优不变得分的目标。仿真实验证实所有得分的名义覆盖率,包括受损网络;在单调变换下,ARC生成相同的集,而插件得分会膨胀;当插件集变为空集时,ARC表现出平滑退化;在well-log基准测试中,ARC将标注的偏移定位到3至5个候选,并通过空集标记不匹配。本文明确指出两个边界:序列相关性会破坏精确性,趋势型替代假设超出分段可交换模型范围。

英文摘要

Conformal changepoint localization turns any score into a confidence set for the changepoint with finite-sample coverage. Coverage is universal; efficiency is not. The oracle score is a likelihood ratio, so practical scores estimate density ratios, and set length deteriorates under heavy tails, skewness, and distribution shift, where no length guarantee applies. We propose ARC (Augmented-Rank Conformalization), a family of scores depending on the data only through within-segment ranks: rank-CUSUM location and scale channels, their fixed combinations, and a lightweight neural score frozen after synthetic training. Every ARC score inherits finite-sample coverage for every frozen weight configuration, including random initialization and mistraining. The main result is an efficiency transfer theorem: the entire ARC confidence set is almost surely invariant under strictly increasing marginal transforms, so the set length distribution depends on the data pair only through its rank structure, and lengths certified once hold verbatim across its monotone orbit, whereas a plug-in score's length changes with every re-expression. Across different rank structures lengths do change, and are reported as such. Classical rank-test theory positions ARC as targeting the optimal invariant score at bounded cost. Simulations confirm nominal coverage for all scores, including sabotaged networks, identical sets under monotone transforms where plug-in scores inflate, and smooth degradation where plug-in sets become vacuous; on the well-log benchmark ARC localizes annotated shifts to three to five candidates and flags misfit by an empty set. Two boundaries are stated rather than hidden: serial dependence destroys exactness, and trend-type alternatives lie outside the piecewise-exchangeable model.

发表机构

  • Beijing University of Technology(北京工业大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑