发表机构
Victoria University of Wellington; University of Canterbury(惠灵顿维多利亚大学; 坎特伯雷大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EMCStereo通过集成三种轻量级注意力模块改进立体匹配,并构建合成树枝数据集VirtualTree,在多个基准上实现高精度细薄结构深度估计。
AI 中文摘要
细薄结构(如树枝)是立体匹配中最困难的情况之一:树枝仅有几个像素宽,背景杂乱,且真实树枝的密集真值几乎无法手工标注。我们做出三项贡献。首先,EMCStereo将三个轻量级注意力模块集成到PSMNet风格的成本体积骨干网络中:对深层语义特征使用高效多尺度注意力(EMA),多尺度融合模块(MSFblock)学习空间金字塔权重而非简单拼接,以及对最终匹配特征使用坐标注意力(CoordAtt)。由于MSFblock将四个金字塔分支压缩为一个,这些模块使网络体积减小2.0%,且仅增加1.7%的推理时间开销。其次,VirtualTree是一个在虚幻引擎5中渲染的合成立体数据集,使用模拟的ZED Mini相机装置,提供5,520对具有精确视差的细树枝图像。第三,一项八路消融研究确立了0.009像素端点误差(EPE)的逐次运行噪声底限。EMCStereo在VirtualTree测试集上达到1.31像素EPE(5.96% D1-all),在SceneFlow上达到1.00像素,在KITTI 2012、KITTI 2015、ETH3D和Middlebury上分别达到0.80、0.73、0.62和3.19像素,深度精度delta_1在92.6%至98.7%之间。与噪声底限相比,在100轮训练预算下,注意力堆叠在精度上保持中性,而MSFblock和CoordAtt在缺少EMA时会造成0.03-0.05像素的误差。
英文摘要
Thin structures such as tree branches are among the hardest cases for stereo matching: a branch is only a few pixels wide, the background is cluttered, and dense ground truth for real branches is nearly impossible to label by hand. We make three contributions. First, EMCStereo integrates three lightweight attention modules into a PSMNet-style cost-volume backbone: Efficient Multi-scale Attention (EMA) on deep semantic features, a Multi-Scale Fusion block (MSFblock) learning spatial pyramid weights instead of concatenating them, and Coordinate Attention (CoordAtt) on final matching features. Because MSFblock collapses four pyramid branches into one, the modules leave the network 2.0% smaller and add only 1.7% inference time overhead. Second, VirtualTree is a synthetic stereo dataset rendered in Unreal Engine 5 with a simulated ZED Mini rig, providing 5,520 pairs with exact disparity for thin branches. Third, an eight-way ablation establishes a run-to-run noise floor of 0.009 px end-point error (EPE). EMCStereo achieves 1.31 px EPE (5.96% D1-all) on the VirtualTree test split, 1.00 px on SceneFlow, and 0.80, 0.73, 0.62, and 3.19 px on KITTI 2012, KITTI 2015, ETH3D, and Middlebury, with depth accuracy delta_1 from 92.6% to 98.7%. Evaluated against the noise floor, the attention stack is accuracy-neutral at a 100-epoch budget, while MSFblock and CoordAtt cost 0.03-0.05 px unless EMA is present.