发表机构
University College London; University of Luxembourg(伦敦大学学院; 卢森堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明重复使用观测数据会虚增预测技能评分,提出保持关联的分配条件及无偏估计方法,并在空气质量数据中验证交互项占比高,不同参考可大幅减少误拒绝。
AI 中文摘要
排名取决于用于比较的观测值。重复使用这些观测值可以在被评估的预测和结果保持不变的情况下,增加预测与结果排名对比之间的关联。我们刻画了在指定总体排名对比(包括预测排名减去基线排名与结果排名减去基线排名的对比)之间保持关联的分配方式。在独立训练的条件下,整条轨迹独立地从共同分布中采样,每条轨迹内部允许无限制的依赖。期望得分分解为其目标部分和地图对之间的显式交互项。当重新分配参考时,我们保持学习到的地图、参考分布和系数行和固定。对于每一对地图,零加权参考重叠是统一保持所有允许地图和分布下关联的必要且充分条件。一个无偏的三轨迹核估计该交互项;独立评估和验证提供了有限样本下界。当所有样本可以重新组合时,完整的U统计量直接估计同一目标。尖锐的共享基线范围(包括并列情况)收紧了两类构造。在北京空气质量档案中,交互项占七种学习预测在经验档案分布下期望共享得分的91.4%至94.5%。另有结果处理类别拟合误差和时间反馈。在匹配的类别对照零实验中,不同的参考将拒绝次数从1000个面板中的402次减少到41次。
英文摘要
Ranks depend on the observations used for comparison. Reusing those observations can add association between forecast and outcome rank contrasts even when the evaluated forecast and outcome stay fixed. We characterize assignments that preserve association between specified population-rank contrasts, including forecast rank minus baseline rank compared with outcome rank minus baseline rank. Conditional on independent training, whole trajectories are sampled independently from a common law, with unrestricted dependence within each trajectory. The expected score separates into its target and an explicit interaction between map pairs. When reassigning references, we keep the learned maps, reference law and coefficient row sums fixed. Zero weighted reference overlap for every map pair is necessary and sufficient for preservation uniformly over permitted maps and laws. An unbiased three-trajectory kernel estimates the interaction; independent evaluation and validation provide finite-sample lower bounds. Complete U-statistics estimate the same target directly when all draws can be recombined. Sharp shared-baseline ranges, including ties, tighten both constructions. In a Beijing air-quality archive, interaction accounts for 91.4% to 94.5% of seven learned forecasts' expected shared scores under the empirical archive law. Separate results address category-fitting error and temporal feedback. In a matched category-control null experiment, distinct references reduce rejections from 402 to 41 out of 1,000 panels.
CommentsAccepted by IEEE ICDM'26