发表机构
Technical University of Munich; Munich Center for Machine Learning; Karlsruhe University of Applied Sciences(慕尼黑工业大学; 慕尼黑机器学习中心; 卡尔斯鲁厄应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究审计7种前馈三维重建骨干网络在13个数据集上的置信度,发现其在非训练条件下不确定性预测偏低,采用幂律修正后偏差可降至1.12倍,并发布审计相关资源。
AI 中文摘要
前馈三维重建模型会输出逐像素的置信度,下游系统将其用作可靠性信号。该置信度被训练为损失权重,而非不确定性幅度,其能否用作误差预测尚未得到测量。我们对7种已发布的骨干网络在13个数据集上进行审计,并从4个属性对置信度评分:其对误差的排序效果、平均水平是否正确、在整个置信度范围内是否有效,以及其区间是否覆盖真实值。结果显示,置信度对误差的排序效果良好,但在非训练条件下使用时,预测的不确定性过低;所有7种模型的中位数偏差达2.4倍,且模型越自信,误差预测偏差越大。我们表明,即使达到损失最优,该现象仍会出现:在训练数据上,已发布模型在自身损失下恢复后,数百次更新内即可达到最优,但对未见过的帧仍过于自信。针对每种骨干网络和数据集,采用含两个常数的幂律可修正预测不确定性的整体幅度,且不改变排序效果。但任何重缩放都无法修正场景层面的问题,我们将其归因于模型缺失跨预测的尺度知识。所有尝试的修正平均接近正确,但仍有三分之一的保留场景落在5点区间外,因为场景缺失的是形状而非偏移。我们发布审计协议、结果,以及每种模型和数据集的拟合常数:在保留目标数据集的情况下拟合时,常数将中位数偏差从2.4倍降至1.35倍;对该数据集的少量标注场景重新拟合后,偏差达1.12倍。
英文摘要
Feed-forward 3D reconstruction models output a per-pixel confidence that is used by downstream systems as an uncertainty signal. The confidence is trained to serve as a weight in the training loss of models. Whether the confidence can be used as an uncertainty magnitude has not been measured. We audit seven backbones on 13 datasets and score the confidence on four properties, i.e., ranking of error, ratio of error to uncertainty on average, slope of this ratio across the confidence range, and coverage of the implied error distribution. Although the confidence ranks error quite well, the uncertainty decoded from the confidence is too small compared to the actual error. The uncertainty has the right size only under the exact training conditions. The median case is off by at least 2.4x across all seven models, while the uncertainty is further off the more confident the model is. Our work shows that the overconfidence appears on unseen scenes even when the model reaches its loss's optimum. As a post-hoc repair we fit a power law on the confidence with two constants per backbone--dataset pair. The repair brings all four audited properties to target at the dataset level, while leaving ranking untouched. Fitted with the target dataset held out, the constants bring the median case from 2.4x off to 1.35x. The repair does not hold below the dataset level, where two-thirds of held-out scenes are still more than five points off in coverage. We attribute what the repair cannot reach to the model, which carries neither the scale of the error nor the shape of its distribution across predictions. We release the audit protocol, its results, and the fitted constants per backbone-dataset pair.