arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00473cs.CVcs.AIcs.CL

CrossProjection:建筑图纸中超越视角变化的几何定位

CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings

Kaho Li, Pengyu Zeng, Yuqin Dai, Jun Yin, Tianjing Feng, Shuai Lu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出CrossProjection方法评估视觉语言模型在异构建筑视图间的几何定位能力,发现模型在自由几何定位上性能脆弱,仅封闭选择成功不代表具备可靠几何定位能力。

中文摘要 AI 辅助

建筑图纸违背了多视图推理背后的常规假设:平面图和剖面图是剖切结果,而立面图是立面投影,因此对应构件的外观变化无法用相机运动来解释。我们引入CrossProjection,这是一种基于锚点的诊断方法,用于判断视觉语言模型是否能在异构建筑视图间保留构件身份并外化几何信息。它通过类别判断、候选选择以及自由点、线、区域定位来评估匹配、配准和几何定位能力。在23套真实图纸及每个模型1954个类别条件下,GPT-5.5得分为82.4%,Qwen3-VL-32B-Instruct为62.2%,GLM-4.5V为57.2%。一项匹配的200个目标研究涵盖了自然图纸和矢量文本被抑制的图纸,以及封闭候选和自由几何输出。有候选支持的性能通常更高,但自由定位仍存在缺陷:在自然图纸上,GPT的点/区域PCK@.05为54-76%,Qwen为8-10%,GLM为14-36%;线端点PCK@.05分别为22%、4%和0%。坐标网格可恢复GPT部分点/区域的精度,但无法改善线的定位。三名接受建筑培训的参与者达到了87.3-93.3%的类别准确率和76-92%的真实区域命中,支持该任务的可行性,而非代表人群层面的人类上限。由于类别族未形成相同项目的匹配-配准对比,且界面控制会改变多重负担,我们不提出机制性主张。可得出的结论更狭隘:封闭选择或标记元素的成功并不意味着可靠的显式几何定位。对于图纸引导的CAD/BIM系统,类别正确性不应被视为无候选空间可靠性的证据。可重复的图纸上锚点、固定分母评分和哈希锁定工件为该差距建立了审计追踪。

英文摘要

Architectural drawings violate the usual assumption behind multi-view reasoning: plans and sections are cuts, while elevations are facade projections, so corresponding components change appearance in ways camera motion cannot explain. We introduce CrossProjection, an anchor-grounded diagnostic of whether vision-language models preserve component identity and externalize geometry across heterogeneous architectural views. It evaluates Matching, Registration, and Geometric Grounding through categorical judgments, candidate selection, and free point, line, and region localization. Across 23 real drawing sets and 1,954 categorical conditions per model, GPT-5.5 scores 82.4%, Qwen3-VL-32B-Instruct 62.2%, and GLM-4.5V 57.2%. A matched 200-target study crosses natural and vector-text-suppressed drawings with closed-candidate and free-geometry outputs. Candidate-supported performance is often higher, but free localization remains fragile: on natural drawings, point/region PCK@.05 is 54-76% for GPT, 8-10% for Qwen, and 14-36% for GLM; line endpoint PCK@.05 is 22%, 4%, and 0%. A coordinate grid recovers some GPT point/region precision but not lines. Three architecture-trained participants reach 87.3-93.3% categorical accuracy and 76-92% GT-region hit, supporting task feasibility rather than a population-level human ceiling. Because the categorical families do not form a same-item Matching-Registration contrast and interface controls alter multiple burdens, we avoid mechanistic claims. The supported conclusion is narrower: closed-choice or marked-element success does not entail reliable explicit geometric grounding. For drawing-guided CAD/BIM systems, categorical correctness should not be treated as evidence of candidate-free spatial reliability. Reusable on-sheet anchors, fixed-denominator scoring, and hash-locked artifacts establish an audit trail for this gap.

发表机构

  • Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
  • University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑