arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20700cs.CV

该病例是否应被适配?预测碎片化控制测试时适配

Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

Lili Wang, Jing Li, Xiaowen Sun, Xiangyu Hu, Zhuangzhuang Gu, Jian Liu, Srihari Nelakuditi, Yan Tong

AI总结:

针对测试时适配中固定步数忽视逐病例差异的问题,提出基于预测碎片化(源模型与适配掩码的分歧几何)的病例级路由器,无需标签或梯度即可预测有害接受面积,在多个医学基准上显著降低损害并保持性能。

AI中文摘要:

情景式测试时适配在每个病例上将冻结的分割器重置为源权重$M_0$,并适配固定的步数。固定视界将队列层面的问题(适配多远)与不可约的逐病例问题(该病例是否应被适配)混为一谈。队列均值掩盖了这一决策:在跨供应商心脏磁共振成像上,适配产生的平均$\u0394$Dice在统计上与零无显著差异,而58.7%的病例个体上变得更差。我们将这种损害量化为有害接受面积(HA),即控制器部署的编辑区域中有害的部分。留出调优比固定视界提供了更强的基线,但其选择的预算在两个主要医学基准上均无法迁移,且没有全局预算能根据病例进行条件化。我们证明预测碎片化——$M_0$与适配掩码$M_k$之间的分歧几何——在决策时无需标签或额外反向传播即可预测HA,在三个基准上表现相当(Spearman $\ ho$ 0.50--0.60),延迟仅为梯度范数的四分之一。基于此构建的病例级路由器在未参与其设计的基准上将HA从0.228降至0.139,设计冻结且仅在该处重新校准截断点。在用于选择设计的的心脏基准上,路由器在匹配的Dice和1.10次部署更新下将HA从0.129降至0.013,对比事后基于评估标签找到的回顾性最优预算,并将58.7%降至20.0%,这是我们所量化的上界。在保留病例未获净收益(如前列腺)的情况下,路由器仍降低HA但牺牲了准确性,这是我们报告的边界。阈值在评估不相交的标记分割上一次拟合;决策不使用标签或梯度。该模板可跨架构和领域移植(nnU-Net$\ o$SegFormer,Cityscapes$\ o$ACDC),坐标、阈值和逐桶操作按领域实例化。

英文摘要:

Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, with an irreducibly per-case one, whether this case should be adapted at all. Cohort means hide that decision: on cross-vendor cardiac MRI the mean $Δ$Dice from adaptation is statistically indistinguishable from zero while 58.7% of cases are individually made worse. We quantify this harm as harmful accepted area (HA), the harmful fraction of the edited area a controller deploys. Held-out tuning gives a stronger baseline than a fixed horizon, but the budget it selects transfers on neither of the two main medical benchmarks, and no global budget can condition on the case. We show that prediction fragmentation---the disagreement geometry between $M_0$ and the adapted mask $M_k$---predicts HA with no labels or extra backward passes at decision time, comparably on three benchmarks (Spearman $ρ$ 0.50--0.60), at a quarter of gradient-norm's latency. A case-level router built on it cuts HA from 0.228 to 0.139 on a benchmark that took no part in its design, with the design frozen and only cut-points recalibrated there. On the cardiac benchmark the design was selected on, the router cuts HA from 0.129 to 0.013 at matched Dice and 1.10 deployed updates, against the retrospective-best budget found post hoc on evaluation labels, and reduces that 58.7% to 20.0%, an upper bound we quantify. Where the retained cases are not net-helped (as on prostate), the router still cuts HA but concedes accuracy, a boundary we report. Thresholds are fit once on a labeled split disjoint from evaluation; decisions use no labels or gradients. The template ports across architecture and domain (nnU-Net$\to$SegFormer, Cityscapes$\to$ACDC) with coordinate, thresholds and per-bucket actions instantiated per domain.

补充信息

↑