arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

iFAN:面向普通掩码 Transformer 的推理感知学习

iFAN: Inference-Aware Learning for Plain Mask Transformers

Fang Li, Yu He, Haoyang Tong, Lichen Ma, Jingling Fu, Wenxiao Fan, Tongxuan Liu, Luohang Liu, Ke Zhang, Junshi Huang

arXiv 2608.03216首次发表:更新:

AI 中文总结

iFAN 是面向普通掩码 Transformer 的推理感知学习框架,通过 APMR 和 CLSD 解决推理与训练不匹配问题,在多个分割数据集上实现性能提升且开销极小。

AI 中文摘要

基于查询的掩码 Transformer 通过最终层查询预测之间的逐像素竞争来组装分割输出,但该推理过程在训练期间未被显式优化。我们发现两个关键不匹配:概率-掩码得分最高的查询不一定能生成最准确的掩码,且最终层解码可能会丢弃中间层的更优预测。为解决这些问题,我们提出了 Inference-Aware Learning(iFAN,推理感知学习),这是一个面向普通掩码 Transformer 的通用训练框架。iFAN 引入 Adjusted Probability-Mask Ranking(APMR,调整后的概率-掩码排名),使查询竞争与预测掩码质量对齐,并抑制高置信度但不准确的竞争者。我们进一步采用 Cross-Layer Self-Distillation(CLSD,跨层自蒸馏),将更强的中间层预测传递到最终层。排名和蒸馏目标仅用于训练,而推理仍保留高效的最终层解码。在 COCO、ADE20K 和 Cityscapes 上的实验表明,其在全景分割、实例分割和语义分割任务中,以及在不同架构、骨干网络规模和输入分辨率下均取得了一致的性能提升。总体而言,iFAN 平均提升了 1.20 PQ、1.30 AP 和 0.63 mIoU,且仅引入可忽略的额外参数、浮点运算量(FLOPs)和推理延迟。

英文摘要

Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of the final layer, yet this inference process is not explicitly optimized during training. We identify two key mismatches: the query with the highest probability-mask score does not necessarily produce the most accurate mask, and final-layer decoding may discard superior predictions from intermediate layers. To address these issues, we propose Inference-Aware Learning (iFAN), a general training framework for plain mask transformers. iFAN introduces Adjusted Probability-Mask Ranking (APMR), which aligns query competition with predicted mask quality and suppresses high-confidence but inaccurate competitors. We further employ Cross-Layer Self-Distillation (CLSD) to transfer stronger intermediate predictions to the final layer. The ranking and distillation objectives are training-only, while inference retains efficient final-layer decoding. Experiments on COCO, ADE20K, and Cityscapes demonstrate consistent improvements across panoptic, instance, and semantic segmentation, as well as across different architectures, backbone scales, and input resolutions. Overall, iFAN improves performance by an average of 1.20 PQ, 1.30 AP, and 0.63 mIoU, with negligible additional parameters, FLOPs and inference latency.

CommentsProject Page https://neesky163.github.io/iFAN/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑