发表机构
CCSE Lab, Beihang University; School of Transportation Science and Engineering, Beihang University; The Hong Kong Polytechnic University(北京航空航天大学CCSE实验室; 北京航空航天大学交通科学与工程学院; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对CAD平面图全景符号识别中现有方法未充分利用文本注释的问题,提出多模态框架TextCAD,设计TACE编码注释语义,引入语义层次对齐框架,实验表明该方法有效提升符号识别性能并达最优。
AI 中文摘要
计算机辅助设计(CAD)平面图包含图形原语和文本注释,为智能设计理解提供互补的几何和语义线索。在CAD分析任务中,随着工业数字化和基于深度学习的自动化需求的增长,全景符号识别变得越来越重要。然而,大多数现有方法主要以原语为中心,未充分利用文本注释。即使是少数文本感知方法也常常只是表面处理注释,没有正确建模CAD注释的复杂语法和层次语义,导致语义丢失和次优的识别性能。为解决这些限制,我们提出TextCAD,一个联合建模图形原语和文本注释用于全景符号识别的多模态框架。具体而言,我们设计了一个类型-属性相关编码器(TACE),通过联合建模注释的类型和属性来明确编码注释中的组合语义。我们还引入了一个带有多级语义过滤(MSF)和原语下采样的语义层次对齐框架,它能在不同语义层次上自适应地将注释语义与图形原语对齐,并实现准确的跨模态语义注入和融合。在真实世界建筑设计数据集上的实验表明,TextCAD有效提高了符号识别性能并取得了当前最优的结果。
英文摘要
Computer-Aided Design (CAD) floor plan drawings contain both graphical primitives and textual annotations, which provide complementary geometric and semantic cues for intelligent design understanding. Among CAD analysis tasks, panoptic symbol spotting has become increasingly important with the growing demand for industrial digitalization and deep learning-based automation. However, most existing methods remain primarily primitive-centric and underexploit textual annotations, despite their critical semantic value. Even the few text-aware approaches often treat annotations only superficially, without properly modeling complex syntax and hierarchical semantics of CAD annotations, which leads to semantic loss and suboptimal spotting performance. To address these limitations, we propose TextCAD, a multimodal framework that jointly models graphical primitives and textual annotations for panoptic symbol spotting. Specifically, we design a Type-Attribute Correlation Encoder (TACE) to explicitly encode the compositional semantics within annotations by jointly modeling their types and attributes. We further introduce a Semantic Hierarchy Alignment framework with Multi-level Semantic Filtering (MSF) and primitive downsampling, which adaptively aligns annotation semantics with graphical primitives at different semantic levels and enables accurate cross-modal semantic injection and fusion. Experiments on real-world building-design datasets show that TextCAD effectively improves symbol spotting performance and achieves state-of-the-art results.