网络应何时输出几何信息,何时检测几何信息?——平面图矢量化中的读出、协调与表示
When Should a Network Emit Geometry, and When Should It Detect It? Readout, Reconciliation, and Representation in Floorplan Vectorization
浏览论文内容
中文总结 AI 辅助
该研究对比了平面图矢量化中两种网络输出方式的性能,发现检测方式在真实扫描数据上表现更优,确定性融合两种输出可提升墙体F1值,还提供了编辑成本指标及相关基准资源。
中文摘要 AI 辅助
用于恢复光栅化平面图的墙体、开口和房间的网络,可通过两种方式输出结果:一是将几何信息作为自回归坐标序列输出,二是在密集的交点和中心线热图上检测几何信息并组装成图。我们在同一训练好的网络上比较这两种读出方式。在真实扫描数据(CubiCasa5K)上,检测方式在所有墙体指标上表现更优(在0.05容差下墙体F1值提升2.7,在0.015容差下提升5.1;配对自举区间不包含零),且使用解码器从未使用过的开口热图可使开口F1值提升2.6倍,无需重新训练。在真实扫描数据中,读出方式的优势随平面图尺寸增大而增强,在小型平面图上则相反;在干净的矢量渲染图上,当渲染风格在训练覆盖范围内时,序列解码的表现优于检测方式5至8个点,而在完全域偏移情况下,若阈值经过校准,检测方式仍领先;墨水密度和平面图尺寸均无法解释这种反转。在匹配的数据和流程下,带协调步骤的以房间为中心的系统与先墙体序列模型达到相当的墙体质量,因此输出表示的重要性通常被低估。另一类方法的先验在输出端有帮助,但在输入端无帮助:两种输出的确定性融合使墙体F1值提升7个点,而用一个模型的输出为另一个模型提供条件则在三种情况下无增益,包括两种真实内容控制情况。我们还提供了编辑成本指标,用于评估修正草稿所需的人工工作量,以及修正后的CubiCasa5K标注,还有ResPlan-FP——一个CC BY 4.0许可的基准,包含16998个平面图,具有固定划分和三条基线轨迹。代码、基准和修正后的标注可在该https URL获取。
英文摘要
A network trained to recover the walls, openings, and rooms of a rasterized floorplan can produce its output in two ways: by emitting the geometry as an autoregressive coordinate sequence, or by detecting it on dense junction and centerline heatmaps and assembling a graph. We compare the two readouts on the same trained network. On real scans (CubiCasa5K) detection is better on every wall measure (+2.7 wall F1 at tolerance 0.05, +5.1 at 0.015; paired bootstrap intervals exclude zero), and reading an opening heatmap the decoder never used raises opening F1 by 2.6x without retraining. Within real scans the readout's advantage grows with plan size and reverses on small plans; on clean vector renders sequence decoding is better by 5 to 8 points where its training covered the render style, while under full domain shift the readout, given calibrated thresholds, stays ahead; neither ink density nor plan size explains the reversal. With matched data and recipe, a room-centric system with a reconciliation step and a wall-first sequence model reach comparable wall quality, so the output representation matters less than is usually assumed. A prior from the other family helps at the output but not at the input: deterministic fusion of the two outputs raises wall F1 by 7 points, whereas conditioning one model on the other's output gives no gain in three forms, including two ground-truth-content controls. We also provide an edit-cost metric that scores a draft by the human work needed to correct it, corrected CubiCasa5K annotations, and ResPlan-FP, a CC BY 4.0 benchmark of 16,998 plans with frozen splits and three baseline tracks. Code, the benchmark, and the corrected annotations are available at https://github.com/Cyprinus12138/fpvec-lab