arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需预训练文本嵌入的主时态取向分类:一种新颖的形态句法库存向量方法

Classifying Dominant Temporal Orientation without Pretrained Text Embeddings: A Novel Morphosyntactic Inventory Vector Approach

Jonathan Cleveland, Peter S. Bearman

arXiv 2609.33121首次发表:更新:

发表机构

Interdisciplinary Center for Innovative Theory and Empirics, Columbia University; Department of Sociology, Columbia University(哥伦比亚大学创新理论与实证跨学科中心; 哥伦比亚大学社会学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种无需预训练嵌入的形态句法库存向量方法,通过词性、依存和将来模式计数编码句子,在1799个复杂句上实现92%的多类时态分类准确率。

AI 中文摘要

计算方法在确定句子的主导时态取向方面一直面临困难。当句子包含多个嵌入从句,且这些从句带有相互竞争的时态和体态信息时,这一困难尤为突出。为解决此问题,我们提出了一种替代方法来识别句子的过去、现在或未来的全局参考区间。我们的方法不使用任何形式的预训练嵌入,而是使用固定长度的库存向量对句子进行编码,这些向量由词性计数、依存关系计数和显式将来模式计数组成。我们将句子的这种库存向量称为“形态句法素”。该方法不使用任何填充、序列模型或大型语言模型。在1,799个标注为过去、现在或未来的句法复杂英语句子上进行评估,结果显示平衡且高准确率的多类分类,总体多类准确率达到92%。

英文摘要

Computational methods have consistently struggled to determine the dominant temporal orientation of a sentence. This difficulty is especially pronounced when a sentence contains multiple embedded clauses with competing tense and aspectual information. To address this difficulty, we propose an alternative approach for identifying a sentence's past, present, or future global reference interval. Our method does not use any form of pretrained embeddings. We rather encode sentences using fixed-length inventory vectors that are comprised of part-of-speech counts, dependency relation counts, and explicit futurate pattern counts. We term this inventory vector of a sentence a "Morphosyntacton". The method does not use any padding, sequence models, or large language models. Evaluation on 1,799 syntactically complex English sentences, annotated as past, present or future, shows balanced and high accuracy multiclass classification, achieving an overall multi-class accuracy of 92%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑