arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20964cs.CLcs.AI

基于带有语义连体相似度评估指标的SAraBERT的阿拉伯语文档抽取式摘要

Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric

Sami Shames El Deen, Mariette Awad

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出增强版AraBERT(SAraBERT)及语义连体相似度评估指标,通过BLEU、ROUGE等验证,证明SAraBERT在阿拉伯语文档抽取式摘要任务中有效。

中文摘要 AI 辅助

本研究提出了SAraBERT,它是AraBERT的增强版本,针对抽取式摘要任务引入了句间Transformer层。为确保SAraBERT生成的摘要能高度覆盖文档的核心思想,我们提出了Semantic Siamese Similarity这一新型评估指标,用于衡量两段文本输入间的相似度水平。我们在Sarabert及已发表相关模型上,使用BLEU、ROUGE和Semantic Siamese相似度进行了验证。模拟结果显示了所提模型的有效性,并为后续研究提供了动力。

英文摘要

In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive summarization tasks. To ensure that the summaries generated by SAraBERT achieve a high coverage of the document's main ideas, we propose Semantic Siamese Similarity, a novel evaluation metric that measures the level of similarity between two text inputs. We validated using BLEU, ROUGE, and Semantic Siamese similarity on Sarabert and published related models. Simulation results showed the effectiveness of our proposed model and motivate follow on research.

发表机构

  • American University of Beirut(贝鲁特美国大学)

机构由 AI 辅助整理,请以论文原文为准。

↑