arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视差有符号:超越零视差平面的立体匹配

Disparity Has a Sign: Stereo Matching Beyond the Zero-Disparity Plane

Jian Shi, Xinge Yang, Chaoyang Wang, Wolfgang Heidrich, Peter Wonka

arXiv 2609.06809首次发表:更新:

发表机构

KAUST(阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视差跨越零值时立体匹配模型失效的问题,提出ZDPShift基准,证明预训练特征已具备负视差能力,仅需训练解码器即可修复,且不损失正值精度。

AI 中文摘要

现代立体匹配模型在视差跨越零值时失效,端点误差(EPE)上升4.6-37倍。然而,从影院3D到VR的立体内容,通常包含位于零视差平面(ZDP)之后的物体,对应负视差。这一盲区级联影响数据集、架构和评估协议,它们都继承了非负几何。校正后的平行相机将ZDP置于无穷远,因此每个有限深度按构造满足$d=fB/z \ge 0$,标准流程中没有任何东西能违反甚至测量负视差。为了测量它,我们提出\textit{ZDPShift}基准,包含来自七个电影摄影师创作的开源电影的$21{,}495$个立体对,每帧在五个零视差平面位置渲染,并带有密集有符号真值。六个最先进的图像和视频立体匹配模型在平面移动后崩溃。在相同场景内容上,FoundationStereo从2.24像素EPE升至75.33像素,每个骨干网络约一半像素超过三像素视差误差。然而,缺失的并非底层匹配能力。使用从SceneFlow合成的监督训练,不添加新数据或参数,误差在有符号范围内保持平坦。仅训练解码器,冻结预训练匹配特征,在全部六个骨干网络上表现相当,EPE在0.2像素内抖动。因此,预训练特征已扩展到它们从未训练过的负值区域,只是输出约定丢弃了它。同时,在KITTI、Middlebury、ETH3D和Sintel上的正值区域精度基本保持。

英文摘要

Modern stereo matching models fail when disparity crosses zero, with end-point error (EPE) rising by 4.6-37$\times$. Yet stereoscopic content, from cinema 3D to VR, routinely contains objects behind the zero-disparity plane (ZDP), corresponding to negative disparities. The blind spot cascades through datasets, architectures, and evaluation protocols, all of which inherit the non-negative geometry. Rectified parallel cameras place ZDP at infinity, so every finite depth yields $d=fB/z \ge 0$ by construction, and nothing within the standard pipeline can violate, or even measure, a negative disparity. To measure it, we propose \textit{ZDPShift}, a benchmark of $21{,}495$ stereo pairs from seven cinematographer-authored open movies, each frame rendered at five zero-disparity-plane positions with dense signed ground truth. Six state-of-the-art image and video stereo matching models collapse once the plane moves. On identical scene content, FoundationStereo goes from $2.24$ px EPE to $75.33$px, with every backbone leaving roughly half of all pixels exceeding a three-pixel disparity error. What is missing, however, is not the underlying matching capability. % The capability itself, however, is already present. Training on supervision synthesized from SceneFlow, which adds no new data or parameters, keeps the error flat across the signed range. Training only the decoder, with the pretrained matching features frozen, performs comparably across all six backbones, with EPE jittering within $0.2$px. Thus, the pretrained features already extend to the negative regime they were never trained on, and only the output convention discarded it. Meanwhile, positive-regime accuracy on KITTI, Middlebury, ETH3D, and Sintel is largely preserved.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑