发表机构
Wichita State University(威奇托州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LOGIC基准通过分离候选选择与传播错误,评估LLM在航空航天电气设计变更影响分析中的意图锚定能力,并探索证据门控与弃权策略以提升安全性。
AI 中文摘要
航空航天电气设计修订可能包含多个真实变更,尽管工程请求可能仅授权其中一部分。因此,传播每一个检测到的差异可能会产生过于宽泛的影响报告。我们提出了LOGIC,一个受控的基准和评估框架,其中本地可部署的语言模型在选定变更通过类型化电气追溯图传播之前,将请求锚定在确定性的候选变更清单中。这种分离使得候选选择错误能够与下游传播错误区分开来。LOGIC包含168个场景,包括144个选择和24个弃权(不执行)案例。我们评估了三个7-8B模型,对比了意图无关、词法和结构化证据方法,并以oracle根作为上界。在96个明确锚定的选择案例中,仅门控结构化证据实现了候选F1为1.0000,而token-词法匹配为0.9677。在12个关系改写案例中,token-词法F1为0.1772,仅门控F1为0.0000,而大型语言模型为0.5000-0.6400。随着候选清单从4个变更增长到64个变更,模型锚定性能下降,而当冻结的选择在约1K到100K节点的图上重放时,受影响元素和类型化路径的准确性保持相对稳定。严格的证据门控抑制了假阳性,但可能移除正确的语义选择。一种探索性的证据为空弃权(不执行)策略将三个模型的严格弃权(不执行)准确性提高到0.6667,并将不安全报告率降低到0.1667,同时将可回答案例覆盖率降低了16.0-27.1个百分点。每个模型的四个冲突请求中有六个仍不安全。这些发现支持在意图无法可靠建立时,将字面证据和语言模型推理与工程审查相结合。
英文摘要
Aerospace electrical-design revisions can contain multiple genuine changes, although an engineering request may authorize only a subset. Propagating every detected difference can therefore produce overly broad impact reports. We present LOGIC, a controlled benchmark and evaluation framework in which locally deployable language models ground a request in a deterministic candidate-change inventory before selected changes are propagated through a typed electrical traceability graph. This separation permits candidate-selection errors to be distinguished from downstream propagation errors. LOGIC contains 168 scenarios, including 144 selection and 24 abstention cases. We evaluate three 7--8B models against intent-agnostic, lexical, and structured-evidence methods, with an oracle-root upper bound. On 96 explicitly anchored selection cases, gate-only structured evidence achieves candidate F1 of 1.0000, compared with 0.9677 for token-lexical matching. On 12 relational-paraphrase cases, token-lexical F1 is 0.1772 and gate-only F1 is 0.0000, compared with 0.5000--0.6400 for the large language models. Model grounding degrades as candidate inventories grow from 4 to 64 changes, while affected-element and typed-path accuracy remain comparatively stable when frozen selections are replayed over graphs of approximately 1K to 100K nodes. Strict evidence gating suppresses false positives but can remove correct semantic selections. An exploratory evidence-empty abstention policy raises strict abstention accuracy to 0.6667 for all three models and reduces unsafe-report rates to 0.1667, while decreasing answerable-case coverage by 16.0--27.1 percentage points. Four of six conflicting requests remain unsafe for each model. These findings support combining literal evidence and language-model reasoning with engineering review when intent cannot be established reliably.
Comments14 pages, 4 figures, 4 tables