发表机构
Nanjing University of Science and Technology; Southeast University(南京理工大学; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出空间提升(SL)方法,通过将2D图像提升到高维空间并用3D网络处理,在保持性能的同时大幅减少参数和推理成本,并支持密集监督与不确定性估计。
AI 中文摘要
我们提出了空间提升(Spatial Lifting, SL),一种用于密集预测任务的新方法。SL通过将标准输入(如2D图像)提升到更高维空间,随后使用为该更高维设计的网络(如3D U-Net)进行处理。与直觉相反,这种维度提升使我们能够在基准任务上获得与传统方法相当的良好性能,同时降低推理成本并大幅减少模型参数数量。SL框架沿提升维度产生内在结构化的输出。这种涌现的结构在训练期间促进密集监督,并在测试时实现基于自一致性的单次前向传播的质量和不确定性估计。空间提升引入了一种简单且通用的建模策略,为视觉中密集预测任务提供了一条通往更高效、更准确、更可靠的深度网络的有前景的路径。
英文摘要
We present Spatial Lifting (SL), a novel methodology for dense prediction tasks. SL operates by lifting standard inputs, such as 2D images, into a higher-dimensional space and subsequently processing them using networks designed for that higher dimension, such as a 3D U-Net. Counterintuitively, this dimensionality lifting allows us to achieve good performance on benchmark tasks compared to conventional approaches, while reducing inference costs and \textbf{drastically lowering the number of model parameters}. The SL framework produces intrinsically structured outputs along the lifted dimension. This emergent structure facilitates dense supervision during training and enables single-forward-pass self-consistency-based quality and uncertainty estimation at test time. Spatial Lifting introduces a simple and general modeling strategy that offers a promising path toward more efficient, accurate, and reliable deep networks for dense prediction tasks in vision.
Comments28 pages 5 figures