arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向话语性文本语义相似度分析的高层文本预处理:框架与实证演示

High-Level Text Preprocessing for Semantic Similarity Analysis of Discursive Texts: A Framework and Empirical Demonstration

Mehmet Murat Albayrakoglu, Mehmet Nafiz Aydin

arXiv 2609.33983首次发表:更新:

发表机构

Işık University; Boğaziçi University(伊斯克大学; 博阿齐奇大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对话语性文本语义相似度因词汇扩散而虚高的问题,提出高层文本预处理框架,含12条规则,经哲学语料实验验证可有效降低相似度分数,并引入语义扩散指数。

AI 中文摘要

语义文本相似度(STS)方法假设文档的词汇内容忠实代表了其主张的内容。这一假设对于在阐述自身立场过程中讨论、比较、批判和情境化其他立场的话语性文档并不成立。结果是语义扩散:文档之间的相似度分数因通过话语性参与而非实质性对齐所获得的词汇而虚高。标准的自然语言处理(NLP)预处理(分词、停用词移除、词干提取、词形还原)无法解决此问题,因为它在词汇层面操作,对所有内容一视同仁,不论其话语功能如何。本文引入了高层文本预处理:一种系统化的、基于规则的干预,在标准预处理流程之前应用,以将每篇文档的实际主张从其话语结构中分离出来。我们提出了12条规则,每条都有明确的理论依据,并在一个百科全书的哲学语料库上演示了其效果:来自斯坦福哲学百科全书的三个条目(美德伦理学、义务论伦理学和后果主义)。一项使用八个基于Transformer的STS模型的三阶段实验表明,预处理降低了所有三个理论对之间的质心余弦相似度分数,24个模型-配对比较中有23个显示出预期的下降,跨模型一致性范围从7-1到8-0。我们引入了语义扩散指数(SDI),一种用于评估文档原始表示与高层预处理表示之间语义重定向的逐文档度量。尽管该框架在哲学文本上进行了演示,但它可能解决一个领域无关的问题,适用于法律文本、政策文件、学术文章以及任何话语性方法引入了文档不认可的立场的词汇的体裁。

英文摘要

Semantic Textual Similarity (STS) methods assume that a document's lexical content faithfully represents what it asserts. This assumption fails for discursive documents that discuss, compare, critique, and contextualize other positions in the process of articulating their own. The result is semantic diffusion: similarity scores between documents are inflated by vocabulary acquired through discursive engagement rather than substantive alignment. Standard Natural Language Processing (NLP) preprocessing (tokenization, stopword removal, stemming, lemmatization) cannot address this problem because it operates at the lexical level, treating all content identically regardless of its discursive function. This paper introduces high-level text preprocessing: a systematic, rule-based intervention applied before the standard preprocessing pipeline to isolate each document's actual claim from its discursive structure. We propose 12 rules, each with an explicit rationale, and demonstrate their effect on an encyclopedic philosophical corpus: three entries from the Stanford Encyclopedia of Philosophy (virtue ethics, deontological ethics, and consequentialism). A three-phase experiment using eight Transformer-based STS models shows that preprocessing reduces centroid cosine similarity scores across all three theory pairs, with 23 of 24 model-pair comparisons showing the expected decrease and cross-model agreement ranging from 7-1 to 8-0. We introduce the semantic diffusion index (SDI), a per-document metric for assessing the semantic reorientation between a document's raw and high-level preprocessed representations. Although the framework is demonstrated using philosophical texts, it potentially addresses a domain-agnostic problem applicable to legal texts, policy documents, academic articles, and any genre in which a discursive approach introduces vocabulary from positions the document does not endorse.

Comments18 pages, 8 tables, 39 references; submitted for publication

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑