arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

句子分割器:用于自监督学习的潜在事实结构揭示

Sentence Splitter: Uncovering Latent Factual Structure for Self-Supervised Learning

Ahmad Pouramini, Mahsa Afsharizadeh

arXiv 2607.19845首次发表:更新:

发表机构

Sirjan University of Technology(锡尔詹理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出句子分割器这一自监督框架,基于T5架构,通过离散分割问题识别句子潜在事实结构,利用自然语言模板训练,可提取前缀-尾部对训练生成模型,实验证明其能提升知识图谱补全和常识问答等下游任务性能。

AI 中文摘要

本文介绍了句子分割器,这是一个基于T5的编码器-解码器架构构建的自监督框架,用于揭示自然语言句子的潜在事实结构。该方法将句子分割表述为离散分割问题,识别描述性前缀(头部)与其事实性完成部分(尾部)之间的语义边界。模型通过概率序列生成来恢复事实性完成部分,而非搜索所有候选边界。为避免人工标注,将符号化的头尾对转化为自然语言模板来训练句子分割器。训练后的分割器用于提取对齐的前缀-尾部对,进而训练生成模型以提出更多合理的完成部分。实验表明该分割器能推广到合成模板之外,且结构感知监督能提升知识图谱补全和常识问答的下游性能。

英文摘要

This paper introduces Sentence Splitter, a self-supervised framework built upon a T5-based encoder--decoder architecture for uncovering the latent factual structure of natural language sentences. The proposed method identifies the semantic boundary between a descriptive prefix (head) and its factual completion (tail) by formulating sentence splitting as a discrete segmentation problem, where a sentence of length $N$ admits $N$ possible split points but only one recovers the intended head--tail structure. Rather than explicitly searching over all candidate boundaries, the model learns to recover the factual completion through probabilistic sequence generation. To eliminate the need for manual annotation, symbolic head--tail pairs are first verbalized into natural-language templates that provide supervision for training the Sentence Splitter. The trained splitter is then applied to raw text to extract aligned prefix--tail pairs, which are subsequently used to train a generative model that proposes additional plausible completions through a lightweight bootstrapping process. This unified pipeline provides a scalable and structure-aware approach to constructing self-supervised training data while bridging symbolic knowledge and natural language. Experiments on both structured and naturally occurring text demonstrate that the proposed splitter generalizes beyond synthetic templates and that the resulting structure-aware supervision consistently improves downstream performance on knowledge graph completion and commonsense question answering, highlighting the effectiveness of recovering latent factual structure for knowledge-centric NLP.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑