发表机构
Data Science Institute, Imperial College London(伦敦帝国理工学院数据科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何在基于归一化的隐式集成中实现可控多样性,核心方法是引入σN-Ens及softmax温度正则化器,主要贡献是在低参数成本下匹配或超越深度集成,能随集成大小扩展且在分布变化下保持校准。
AI 中文摘要
深度集成在深度学习中能提供最可靠的不确定性估计,但成本随成员数量线性增长。隐式集成通过共享单个主干降低成本,然而成员多样性在训练中无法塑造。我们引入了σN-Ens,一种基于归一化的隐式集成,将每个成员视为多任务架构中的一个任务,并通过sigmoid有界缩放器调制共享主干。还引入了softmax温度正则化器,塑造成员间的共享平衡水平并追踪准确性校准前沿。只复制归一化层,该机制可应用于卷积和变压器主干,也能通过微调预训练模型。我们将这种集成表达的认知不确定性视为调制不确定性,并解释了其校准在输入损坏下成立的原因以及其分布外检测较弱的原因。在CIFAR-10/100、ImageNet和SST-2上对ResNets和变压器评估了我们的方法。σN-Ens以参数成本的一小部分匹配或优于深度集成,在分区方法失效时随集成大小扩展,并在分布变化下保持校准。
英文摘要
Deep ensembles provide the most reliable uncertainty estimates in deep learning, but their cost grows linearly with the number of members. Implicit ensembles lower this cost by sharing a single backbone across members. Member diversity is a primary determinant of ensemble quality, yet no implicit ensemble can shape it during training; existing methods fix it at initialisation or build it into the architecture. We introduce $σ$N-Ens, a normalisation-based implicit ensemble that treats each member as a task in a multi-task architecture and modulates the shared backbone through sigmoid-bounded scalers. We also introduce a softmax-temperature regulariser, which shapes the equilibrium level of sharing between members and traces the accuracy-calibration frontier. Because only normalisation layers are replicated, the mechanism can wrap convolutional and transformer backbones alike, also allowing pretrained models to be adapted through a short fine-tune. We frame the epistemic uncertainty such an ensemble expresses as modulation uncertainty, and explain why its calibration holds under input corruption, and why its out-of-distribution detection is weaker. Our method is evaluated across ResNets and transformers on CIFAR-10/100, ImageNet and SST-2. $σ$N-Ens matches or outperforms deep ensembles at a fraction of their parameter cost, scales with ensemble size where partitioning methods collapse, and maintains calibration under distribution shift.