arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10378cs.CL

使用大语言模型进行爱沙尼亚语文档级文本简化

Document-Level Text Simplification in Estonian Using Large Language Models

Meeri-Ly Muru, Eduard Barbu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估五种多语言大语言模型在爱沙尼亚语文档级简化中的表现,提出三种提示策略和综合评估框架,发现Gemini-2.0和LLaMA-3.3效果最佳,并贡献了新的连贯性指标和公开资源。

中文摘要 AI 辅助

文档级文本简化涉及超越句子内部编辑的转换,处理话语连贯性、指代消解和跨段落一致性。尽管在高资源语言的句子级简化方面取得了进展,但在形态丰富、低资源语言(如爱沙尼亚语)中的文档级简化仍 largely 未被探索。本研究对五种最先进的多语言大语言模型(LLMs)在爱沙尼亚语文档级简化中进行了全面评估。研究考察了三种提示策略:单次生成、基于流水线的模块化智能体以及指南增强的流水线。评估框架整合了评估可读性、语义保留和话语连贯性的自动指标,以及结构化的手动标注协议。研究结果表明,Gemini-2.0 和 LLaMA-3.3 生成的输出具有接近母语的流畅性和强大的意义保留能力,而其他模型则表现出显著的语法和语义局限性。这项工作贡献了新颖的文档级连贯性指标、基于证据的提示策略以及用于可复现性的公开可用资源。

英文摘要

Document-level text simplification involves transformations that go beyond sentence-internal edits, addressing discourse coherence, anaphora resolution, and cross-paragraph consistency. Despite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored. This study presents a comprehensive evaluation of five state-of-the-art multilingual large language models (LLMs) for document-level simplification in Estonian. Three prompting strategies are examined: single-pass generation, pipeline-based modular agents, and guideline-augmented pipelines. The evaluation framework integrates automatic metrics assessing readability, semantic preservation, and discourse coherence, alongside a structured manual annotation protocol. The findings indicate that Gemini-2.0 and LLaMA-3.3 produce outputs with near-native fluency and strong meaning preservation, whereas other models display notable grammatical and semantic limitations. This work contributes novel document-level coherence metrics, evidence-based prompting strategies, and publicly available resources for reproducibility.

发表机构

  • National Library of Estonia(爱沙尼亚国家图书馆)
  • University of Tartu(塔尔图大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑