arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规范文本的论元感知语义对齐:一种基于图尔敏的神经符号方法

Argument-Aware Semantic Alignment of Normative Texts: A Toulmin-Based Neuro-Symbolic Approach

William Schroeder

arXiv 2608.29529首次发表:更新:

发表机构

Purdue University; Cleantech Software(普渡大学; 清洁技术软件公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对专业规范文本语义对齐难题,提出结合神经文本表示与图尔敏特征的神经符号方法,在网络安全标准映射基准上验证论元结构特征可提升对齐效果。

AI 中文摘要

当等效要求使用不同术语、句法和抽象程度时,专业规范文本之间的语义对齐颇具挑战性。词汇重叠、分布嵌入和语义相似度虽能捕捉主题相关性,但常忽略支撑、限定和论证规范主张的论元结构。本文探究显式论元结构是否能为要求对齐补充神经语义信息。我们将跨标准控制映射视为论元感知语义对齐,构建结合神经文本表示与图尔敏(Toulmin)特征的神经符号管道。大语言模型(LLM)的显式化步骤可识别主张(claim)、依据(ground)、正当理由(warrant)、限定词(qualifier)和支持(backing)并重构省略三段论(enthymeme)。这些内容通过论元感知相似度和结构特征输入对齐模型。在NERC-CIP到NIST-CSF的映射基准上,论元衍生特征较神经符号语义基线提升了对齐效果。特征选择显示与正当理由相关的特征信号尤为强烈,表明主张与其支撑推理间的联系无法仅通过传统相似度捕捉。紧凑的主张-依据-正当理由子集仍与完整图尔敏特征集具有竞争力。结果初步证明论元结构是对齐专业规范文本的有用中间表示。网络安全标准被用作受控测试平台,而非领域无关泛化的证明。LLM显式化生成的论元图还可支持后续针对规范文本的检索、推理和解释工作。

英文摘要

Semantic alignment between specialized normative texts is challenging when equivalent requirements use different terms, syntax, and levels of abstraction. Lexical overlap, distributional embeddings, and semantic similarity capture topical relatedness but often miss the argumentative structure by which normative claims are supported, qualified, and justified. This paper asks whether explicit argument structure adds information complementary to neural semantics for aligning requirements. We treat cross-standard control mapping as argument-aware semantic alignment and build a neuro-symbolic pipeline that combines neural text representations with Toulmin features. An LLM explicitation step identifies claims, grounds, warrants, qualifiers, and backing and reconstructs enthymemes. These feed an alignment model via argument-aware similarity and structural features. On a NERC-CIP to NIST-CSF mapping benchmark, argument-derived features improve alignment over a neuro-symbolic semantic baseline. Feature selection shows especially strong signal from warrant-related features, indicating that the link between a claim and its supporting reasoning is not captured by conventional similarity alone. A compact claim--grounds--warrant subset remains competitive with the full Toulmin feature set. The results give preliminary evidence that argument structure is a useful intermediate representation for aligning specialized normative texts. Cybersecurity standards are used as a controlled testbed, not as proof of domain-independent generalization. The argument graphs produced by LLM explicitation may also support later work on retrieval, reasoning, and explanation over normative text.

Comments11 tables, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑