发表机构
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究小型本地部署语言模型对LEED文档筛选及符号组件作用,引入神经符号管道,通过对齐PDF、检索证据、语言模型验证和数值检查器检查来验证合规性,实验表明特定模型在任务中表现较好,管道及基线提供了参考。
AI 中文摘要
LEED v4.1 BD+C认证仍是一个文档密集型过程,要求审核人员阅读数百页项目证据并手动应用特定信用的阈值逻辑。本文研究小型本地部署语言模型能否对LEED文档进行有意义的筛选,以及确定性符号组件应如何分担此项工作。引入了一个神经符号管道,将项目PDF与LEED信用部分对齐,通过信用感知关键词签名检索证据,使用本地托管的40亿参数语言模型验证合规性,并对定量阈值应用特定于LEED的数值检查器。对四所大学建筑(484个PDF,153个信用级决策)的实验表明,40亿参数模型(gemma3:4b)是最强的纯文本核心验证器,准确率达67.3%,在此任务中优于更大的80亿参数模型(llama3.1:8b)。确定性数值检查器纠正了关键定量信用上算术错误,将EA-p2的准确率从50%提高到100%,并在可靠提取所需值时改善了其他几个信用。同时,完整的神经符号配置总体准确率为61.6%,因提取失败和在定性类别上的保守行为落后于最佳纯文本基线。系统消融表明,添加低分辨率绘图图像(150-300 dpi)会持续降低准确率,提示有效性取决于建筑物的实际通过率:规则提示在文档丰富的项目上表现最佳,而思维链提示在文档较少的项目上表现最佳。在针对原始项目文档的LEED v4.1 BD+C合规验证的特定范围内,此管道及其基线为准确率和失败模式提供了一个初步可重复的参考点。
英文摘要
LEED v4.1 BD+C certification remains a document-intensive process that requires reviewers to read hundreds of pages of project evidence and apply credit-specific threshold logic by hand. This paper investigates whether small, locally deployed language models can perform meaningful screening of LEED documentation and how deterministic symbolic components should share that work. A neuro-symbolic pipeline is introduced that aligns project PDFs to LEED credit sections, retrieves evidence with credit-aware keyword signatures, verifies compliance with a locally hosted 4-billion-parameter language model, and applies a LEED-specific numeric checker to quantitative thresholds. Experiments on four university buildings (484 PDFs, 153 credit-level decisions) show that a 4-billion-parameter model (gemma3:4b) is the strongest text-only core verifier, achieving 67.3% accuracy and outperforming a larger 8-billion-parameter model (llama3.1:8b) in this task. The deterministic numeric checker corrects arithmetic errors on key quantitative credits, moving EA-p2 from 50% to 100% accuracy and improving several other credits when required values are reliably extracted. At the same time, the full neuro-symbolic configuration achieves 61.6% overall accuracy, trailing the best text-only baseline due to extraction failures and conservative behavior on qualitative categories. Systematic ablations show that adding low-resolution drawing images (150-300 dpi) consistently reduces accuracy, and that prompt effectiveness depends on the building's ground-truth PASS rate: rubric prompts perform best on documentation-rich projects, while chain-of-thought prompts perform best on documentation-lean projects. Within the specific scope of LEED v4.1 BD+C compliance verification over raw project documentation, this pipeline and its baselines provide an initial reproducible reference point for both accuracy and failure modes.