发表机构
The University of Adelaide; Australian National University; Australian Institute for Machine Learning (AIML)(阿德莱德大学; 澳大利亚国立大学; 澳大利亚机器学习研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现冻结DETR分割模型存在输出选择瓶颈,提出仅基于缓存输出的轻量选择器HYDRA,在不增加计算下显著提升多个分割器的性能,揭示模型已计算但未暴露的潜在掩码知识。
AI 中文摘要
现代分割器常常在昂贵的计算已经完成后才失败:一个有用的掩码存在于模型的查询条件候选集中,但部署的选择规则并未将其暴露出来。我们研究了冻结的DETR系列模型中的这种输出选择瓶颈。一个仅使用真实标签的预言机首先显示,在已计算的掩码候选中存在大量的隐藏余量。这引出了一个简单的问题:我们如何更好地利用分割器已经计算但未暴露的掩码?随后,我们探究是否可以在不添加查询、不生成新掩码、不重新运行骨干网络或更新权重的情况下恢复该余量。HYDRA是一个仅在缓存的冻结输出上训练的小型选择器。在推理时,它根据一个显式的保留基线选项对缓存的候选进行评分,并且仅在保留的校准边际表明所选候选足够好时才采取行动。HYDRA在训练集缓存上训练并在保留数据上校准,在ADE20k和COCO上将Mask2Former、MaskDINO和OneFormer的数据集mIoU分数最多提高了+7.41个百分点,并在八个域中将SAM 3的类宏提示IoU平均提高了+9.4个百分点,同时通过校准保留了有用的预测。配对的LoRA对照实验表明,轻量级权重适应并不能消除瓶颈:暴露的预测通常是持平或更差,而对适应后的候选进行路由仍然可以恢复准确性。最后,我们将该效果与二分匹配下的查询特化联系起来,并在受控的TinyDETR研究中验证了这一点。这些结果表明,冻结的分割器不仅应根据其暴露的掩码来评估,还应根据其抑制的有用候选来评估。
英文摘要
Modern segmenters often fail after the expensive computation has already been done: a useful mask is present among the model's query-conditioned candidates, but the deployed selection rule does not expose it. We study this output-selection bottleneck in frozen DETR-family models. A ground-truth-only oracle first shows substantial hidden headroom in already-computed mask proposals. This raises a simple question: How can we better use the masks a segmenter has already computed but does not expose? We then ask whether that headroom can be recovered without adding queries, generating new masks, rerunning the backbone, or updating weights. HYDRA is a small selector trained only on cached frozen outputs. At inference time, it scores the cached candidates against an explicit keep-baseline option and acts only when a held-out calibrated margin indicates the selected candidate is sufficiently better. Trained on training-split caches and calibrated on held-out data, HYDRA improves Mask2Former, MaskDINO, and OneFormer by up to +7.41 dataset mIoU points on ADE20k and COCO, and improves SAM 3 by +9.4 class-macro prompt-IoU points on average across eight domains while preserving useful predictions through calibration. Paired LoRA controls show that lightweight weight adaptation does not remove the bottleneck: exposed predictions are often flat or worse, while routing over the adapted candidates still recovers accuracy. Finally, we connect the effect to query specialization under bipartite matching and verify it in a controlled TinyDETR study. These results show that frozen segmenters should be evaluated not only by the masks they expose, but also by the useful candidates they suppress.