发表机构
Korea Electronics Technology Institute(韩国电子技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出纵向协议测量窄域微调对检测器覆盖预训练词汇外对象的影响,发现覆盖度下降而域内精度上升,并提出无需训练的修复方法,混合预训练状态可恢复覆盖且域内精度损失极小。
AI 中文摘要
一个在广泛语料上预训练的检测器在窄域上进行微调,其域内准确率提升,然后被部署。我们询问在此期间,其对词汇表从未命名的对象的覆盖发生了什么,这些对象在障碍物检测和检查中带来风险。没有域内测试集包含此类对象的示例。我们提出一个纵向协议:一个预训练检查点与其自身的微调后代进行比较。它跟踪保留的Top-K提议覆盖度C_τ:预训练覆盖但词汇表省略的类别中,检测器Top-K区域仍覆盖的框的比例。该量来自开放世界提议文献;纵向解读则是新的。C_τ在域内准确率上升时下降,在四个架构和三个域上,对于面积超过1024像素平方的框,下降幅度为5.12到63.35个百分点。没有域内指标能识别这种下降,检测平均精度也不能,因为它对漏检和误命名框同等扣分。在同时评估两者的一个架构上,适应损失了AP的87%,而覆盖度仅损失五分之一,并且在其冻结阶梯的所有六个深度上,每次运行都是命名先丢失。所破坏的是结构性的:三个共享预训练运行的架构在哪些类别丢失覆盖上达成一致,而那些模型从未学过的类别不会丢失任何覆盖。随后进行修复,无需训练:将预训练状态的四分之一混合回去,包括归一化统计量,在扫描的每个单元上提高覆盖度,最多损失2.47个百分点的域内准确率。看到这一点需要额外一次评估遍历。
英文摘要
A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhile to its coverage of objects the vocabulary never names, which in obstacle detection and inspection carry the risk. No in-domain test set holds an example of one. We give a longitudinal protocol: one pretrained checkpoint against its own fine-tuned descendants. It tracks held-out top-$K$ proposal coverage $C_τ$: of categories pretraining covered and the vocabulary omits, the share of boxes a detector's top $K$ regions still cover. The quantity is the open-world proposal literature's; the longitudinal reading is not. $C_τ$ falls while in-domain accuracy rises, on four architectures and three domains, by $5.12$ to $63.35$ points on boxes above $1024$ px$^2$. No in-domain number identifies the fall, and neither does detection average precision, which charges a missed and a misnamed box alike. On the one architecture scoring both, adaptation costs $87\%$ of the AP against a fifth of the coverage, and the naming goes first at all six depths of its freeze ladder, every run. What breaks is structured: three architectures sharing no pretraining run agree on which categories lose coverage, and those a model never learned do not lose any. A repair follows and needs no training: mixing a quarter of the pretrained state back, normalisation statistics included, raises coverage on every cell swept for at most $2.47$ points of in-domain accuracy. Seeing it costs one extra evaluation pass.
Comments25 pages, 3 figures. Supplementary material (69 pages) is included as an ancillary file. Submitted to the International Journal of Computer Vision