动态集成何时奏效?诊断分布偏移下回归任务中的区域级增益
When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift
浏览论文内容
中文总结 AI 辅助
针对分布偏移下动态回归集成增益难预判的问题,提出$\text{D}_{\rm CF5}$指标通过少量目标域探测数据估计区域级增益,结合门控选择器实现安全部署,并发布了可复现评估基准OpenRegShift。
中文摘要 AI 辅助
回归模型池的输入依赖型(“动态”)组合是否优于最优静态混合,取决于偏移情况,且在部署前通常无法得知。少量带标注的目标域探测数据能否告诉我们,在输入空间的不同区域重新分配模型信任度何时会带来收益?我们通过$\boldsymbol{\text{DCF5}}$(交叉拟合5折区域增益估计量)回答了这一问题,该指标从探测数据中估计区域级凸组合相对于最优静态凸混合的交叉拟合增益,即逐区域决定信任哪个模型的可实现价值。\n在包含12个数据集-偏移对(空间、时间、域、特征聚类类型)的固定测试套件中,$\text{D}_{\rm CF5}$预测实际区域级测试增益的数据集层面Spearman相关系数为$+0.98$(95%置信区间$[+0.83, +1.00]$;$p=5\times10^{-5}$),其中有两个案例推翻了预注册的预期。在16对样本的敏感性分析中该相关性依然成立(Spearman系数$+0.83$),而其他备选探测诊断方法的相关系数最高仅为$+0.66$。这种差异明确了区域信任重分配的作用:区域凸增益的相关系数为$+0.98$,而经仿射校正的平滑协变量依赖堆叠的相关系数仅为$+0.01$。\n受控生成器实验表明,动态增益源于偏移异质性与局部能力的相互作用,随偏移程度加剧而提升,且在测试网格中,当探测标注数量在128到256之间时,动态增益即可实现。“探测验证集成选择器”在静态仿射堆叠器和动态实现器之间做选择,仅当留出集的下置信界超过静态凸基线时才部署候选模型。在预注册的前瞻性批次实验中,它在全部12次运行中均达到或优于基线;两次部署分别将测试风险降低了11%和16%,而门控机制拒绝了一个候选模型——若未加门控部署该模型,产生的损失将超过静态损失的$30\times$。我们还发布了OpenRegShift,这是一个用于分布偏移下回归集成的可复现评估框架。
英文摘要
Whether input-dependent ("dynamic") combination of a regression model pool beats the best static blend depends on the shift and is rarely known before deployment. Can a small labeled target-domain probe tell us when reallocating trust across regions of the input space will pay off? We answer this with $\widehat{D}_{\mathrm{CF5}}$, which estimates from the probe the cross-fitted gain of the regionwise convex combination over the best static convex blend: the realizable value of deciding, region by region, whom to trust. Across a frozen suite of 12 dataset-shift pairs (spatial, temporal, domain, feature-cluster), $\widehat{D}_{\mathrm{CF5}}$ predicts realized regionwise test gains with dataset-level Spearman $+0.98$ (95% CI $[+0.83, +1.00]$; $p=5\times10^{-5}$), including two cases overturning preregistered expectations. The relationship holds in a 16-pair sensitivity analysis (Spearman $+0.83$), whereas alternative probe diagnostics reach at most $+0.66$. This contrast isolates regional trust reallocation: correlation is $+0.98$ for regionwise-convex gain, but $+0.01$ for smooth covariate-dependent stacking after affine correction. A controlled generator shows dynamic gains arise from the interaction of shift heterogeneity and local competence, increase with shift severity, and become realizable between 128 and 256 probe labels in the tested grid. The Probe-Validated Ensemble Selector chooses among a static affine stacker and dynamic realizers, deploying a candidate only when a held-out lower confidence bound clears the static-convex floor. In a preregistered prospective batch, it matched or improved the floor in all 12 runs; two deployments reduced test risk by 11% and 16%, while the gate rejected a candidate whose un-gated deployment incurred $>30\times$ the static loss. We release OpenRegShift, a reproducible evaluation harness for regression ensembles under distribution shift.