arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Shaer:基于格律子形式和语义条件的受控阿拉伯诗歌生成

Shaer: Controlled Arabic Poetry Generation with Meter Subform and Semantic Conditioning

Ahmad Abbas, Tamara Fakih, Nour Fakih, Ammar Mohanna

arXiv 2610.09756首次发表:更新:

发表机构

American University of Beirut(贝鲁特美国大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Shaer提出受控古典阿拉伯诗歌生成框架,联合语义、格律子形式和长度条件,基于增强语料微调Yehia-7B,显著提升格律与长度控制精度。

AI 中文摘要

古典阿拉伯诗歌生成需要同时满足语义、语言和细粒度韵律约束。现有系统通常控制宽泛的诗歌属性,但未联合建模语义意图、格律子形式和诗歌长度。我们提出Shaer,一个受控古典阿拉伯诗歌生成框架,联合基于自然语言描述、格律子形式和目标半句数进行条件生成。为支持该任务,我们构建了一个包含116,032首古典阿拉伯诗歌的增强语料库,源自Ashaar,包含归一化的格律子形式标签和自动生成并验证的语义描述。随后,我们使用基于QLoRA的监督微调,采用仅补全目标,对Yehia-7B进行适配。我们的评估结合了基础格律符合度、请求子形式遵循度和长度控制的自动评估,以及三个LLM评审、盲法人工评估和记忆分析。Shaer实现了95.17%的基础格律准确率、91.75%的诗歌级格律子形式准确率和83.40%的精确计数准确率。相对于未微调的基础模型,这些结果分别提升了68.68、57.77和38.93个百分点;Shaer还在所有评估系统中取得了最高的基础格律准确率。多LLM评估和对排名靠前输出的盲法人工评估进一步表明其语义和文学质量具有竞争力。最后,对所有3,481个测试生成的分析未发现训练语料库或配对源诗的精确副本。代码、模型和数据集公开可用。

英文摘要

Classical Arabic poetry generation requires simultaneously satisfying semantic, linguistic, and fine-grained prosodic constraints. Existing systems typically control broad poetic attributes but do not jointly model semantic intent, meter subform, and poem length. We present Shaer, a controllable Classical Arabic poetry generation framework jointly conditioned on natural-language descriptions, meter subforms, and target hemistich counts. To support this task, we construct an enriched corpus of 116,032 classical Arabic poems derived from Ashaar, containing normalized meter-subform labels and automatically generated, validated semantic descriptions. We then adapt Yehia-7B using QLoRA-based supervised fine-tuning with a completion-only objective. Our evaluation combines automatic assessment of base-meter conformity, requested-subform adherence, and length control with three LLM judges, blinded human evaluation, and memorization analysis. Shaer achieves 95.17% base-meter accuracy, 91.75% poem-level meter-subform accuracy, and 83.40% exact count accuracy. Relative to its untuned foundation model, these results represent gains of 68.68, 57.77, and 38.93 percentage points, respectively; Shaer also attains the highest base-meter accuracy among all evaluated systems. Multi-LLM evaluation and a blinded human assessment of top-ranked outputs further indicate competitive semantic and literary quality. Finally, analysis of all 3,481 test generations finds no exact copies from the training corpus or paired source poems. Code, models, and datasets are publicly available.

Comments22 pages, 7 figures. Code: https://github.com/AhmaddAbbass/Shaer ; models and datasets: https://huggingface.co/Shaer-AI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑