AI 中文总结
本文提出无领域知识的元认知层,通过标签向量池学习几何规则,将多个ViT检测器的融合建模为溯因问题,在航拍基准测试中对抗标签翻转攻击时性能优于多数投票基线。
AI 中文摘要
将预训练感知模型部署到新环境时,其准确率会因分布偏移而下降,仅将这些模型组合起来无法恢复准确率:多数投票等组合器会以牺牲召回率为代价换取准确率,且对协同失效情况较为脆弱。现有的元认知方法会学习用于标记模型错误的逻辑规则,但依赖人工编写的领域知识线索(如目标大小先验、分割掩码),这些线索无法迁移到真正新颖的场景中。本文展示,通过利用向量空间几何,无需任何领域知识即可学习该元认知层:基于每个模型自身训练嵌入构建的单模型标签向量池(Label Vector Pools, LVP),可根据检测结果相对于训练确定的原型的几何结构生成错误检测规则,在测试集上的F1值与领域知识规则的差距不超过0.002。由于该方法仍属于神经符号方法,这些几何规则共享单一逻辑框架,在有领域知识时仍可补充。本文将多个不完美的基于ViT的检测器的融合问题建模为基于一致性的溯因问题,在测试时通过精确整数规划(Integer Program, IP)和多项式时间启发式算法求解。在包含15个天气偏移测试集和6个ViT检测器的航拍基准测试中,本文提出的无领域知识层在干净数据上的表现与最强的多数投票变体相当(F1值差距在0.005以内),且与所有多数投票基线不同,它在协同标签翻转攻击下能保持性能:在90%的翻转率下,其平均F1值为0.42,而MV-Plurality的平均F1值为0.35(相对提升22%),且当翻转率超过0.4时,在每个测试集上都达到最高F1值。
英文摘要
Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not recover it: combiners such as majority voting trade recall for precision and are brittle to coordinated failures. Prior metacognitive methods learn logical rules that flag a model's errors, but rely on hand-authored domain-knowledge cues (object-size priors, segmentation masks) that do not transfer to genuinely novel scenes. We show that this metacognitive layer can be learned without any domain knowledge by exploiting vector-space geometry: per-model Label Vector Pools (LVP), built from each model's own training embeddings, yield error-detection rules from the geometry of detections relative to training-determined prototypes, reaching parity with domain-knowledge rules to within $0.002$ every F1 on test set. Because the approach remains neurosymbolic, these geometric rules share a single logical framework and can still be complemented by domain knowledge when available. We frame the fusion of multiple imperfect ViT-based detectors as a consistency-based abduction problem solved at test time by an exact Integer Program (IP) and a polynomial-time heuristic. On an aerial-imagery benchmark of 15 weather-shifted test sets and six ViT detectors, our domain-knowledge-free layer matches the strongest majority-vote variant on clean data (within $0.005$ F1) and, unlike every majority-vote baseline, retains its performance under a coordinated label-flipping attack: at a $90\%$ flip rate it averages $0.42$ F1 versus $0.35$ for MV-Plurality (a $22\%$ relative gain) and attains the highest F1 on \emph{every} test set once the flip rate exceeds $0.4$