AI 中文总结
本文提出Depth Any Seen方法,利用多伯努利深度集和ExactMB目标联合估计单目图像中可见表面的存在性与度量深度,显著减少过度预测,提升分层深度估计准确性。
AI 中文摘要
当一条射线上可见多个表面时,从单张图像恢复可见的3D结构需要联合估计它们的存在性和度量深度。深度任意可见(Depth Any Seen)将这些表面表示为图像条件下的多伯努利深度集,其每个分量贡献一个深度或保持缺失。其无辅助的精确多伯努利目标(ExactMB)通过边缘化到完整、不同目标的一对一分配来学习深度和存在性。我们的分析表明,匹配期望计数可能留下分量-表面分配未解决。我们扩展了真实和合成分层深度基准,以评估深度准确性、恢复的支持度和过度预测。与深度堆叠相比,ExactMB在LD-Real上将过度预测相对减少88.2%,在MD-3K上减少80.5%,同时保留了大部分序数准确性,在LD-Syn上具有可比较的条件度量深度误差。进一步的消融研究表明,有序分配比边缘化提高了深度准确召回率和精确率,而计数正则化配置在较低召回率下实现了比有序分配更高的更深排名精确率。我们的代码将公开发布。
英文摘要
When several surfaces are visible along a ray, recovering visible 3D structure from one image requires jointly estimating their presence and metric depth. Depth Any Seen represents these surfaces as image-conditioned multi-Bernoulli depth sets, whose components each contribute one depth or remain absent. Its auxiliary-free Exact Multi-Bernoulli objective (ExactMB) learns depth and presence by marginalizing one-to-one assignments to complete, distinct targets. Our analysis shows that matching expected count can leave component-surface assignment unresolved. We extend real and synthetic layered-depth benchmarks to evaluate depth accuracy, recovered support, and overprediction. Compared to depth stacking, ExactMB reduces overprediction by a relative 88.2% on LD-Real and 80.5% on MD-3K while retaining most ordinal accuracy, with comparable conditional metric-depth error on LD-Syn. Further ablation studies show that ordered assignment improves depth-accurate recall and precision over marginalization, whereas the count-regularized configuration achieves higher deeper-rank precision than ordered assignment at lower recall. Our code will be publicly released.
Comments54 pages, 31 figures, including appendix. Video demo: https://youtu.be/D8TIzhjq-2Y