arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越诗歌:大型语言模型能否生成古典阿拉伯语玛卡梅?

Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?

AbdulRahman A. Morsy, Aya Zirikly

arXiv 2609.28245首次发表:更新:

发表机构

George Washington University; Johns Hopkins University(乔治华盛顿大学; 约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次系统评估大型语言模型生成古典阿拉伯语玛卡梅的能力,比较五种模型在不同提示策略下的表现,发现少样本提示提升押韵密度,而零样本提示总体得分最高,揭示了模型在文体对齐上的差异。

AI 中文摘要

大型语言模型(LLMs)在创意文本生成方面表现出强大的性能,但其生成具有文化根基和文体约束的文学形式的能力仍未得到充分探索。先前的工作主要集中在现代语言变体和诗歌上,而诸如玛卡梅(maqama)之类的古典散文传统在很大程度上仍未得到研究。玛卡梅是一种古典文学体裁,其特征是押韵散文(saj)、密集的修辞修饰和片段式叙事结构,使其成为评估LLMs能否超越表面流畅性、迈向更深层文学能力的具有挑战性的测试平台。在本文中,我们提出了首个针对LLMs生成玛卡梅的受控评估研究,在零样本、少样本和基于规则的提示下比较了五个模型,并通过人工标注和LLM作为评判者的框架,在修辞丰富性、saj密度、结构连贯性和文体真实性等维度上评估输出。我们的结果表明,提示策略在文体质量中起着重要作用:少样本提示最一致地提高了saj密度,而其对修辞和连贯性的影响因模型而异,最强的模型(GPT-4o和GPT-5.4-mini)在这些维度上从基于规则的提示中获益最多,尽管零样本提示在所有五个模型中产生了最高的总体得分。我们进一步观察到模型在与阿拉伯语玛卡梅惯例的文体对齐方面存在系统性差异,并通过第二个独立的LLM评判者、配对统计显著性检验和saj的非LLM代理测量来证实我们的发现。

英文摘要

Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored. Prior work has focused largely on modern language varieties and poetry, while classical prose traditions such as maqama remain largely unstudied. The maqama is a classical literary genre characterized by rhymed prose (saj), dense rhetorical ornamentation, and episodic narrative structure, making it a challenging testbed for evaluating whether LLMs can move beyond surface fluency toward deeper literary competence. In this paper, we present the first controlled evaluation study of maqama generation with LLMs, comparing five models under zero-shot, few-shot, and rule-based prompting, and evaluating outputs through both human annotation and an LLM-as-a-judge framework across dimensions such as rhetorical richness, saj density, structural coherence, and stylistic authenticity. Our results show that prompting strategy plays a strong role in stylistic quality: few-shot prompting most consistently improves saj density, while its effects on rhetoric and coherence vary by model, with the strongest models (GPT-4o and GPT-5.4-mini) benefiting most from rule-based prompting on these dimensions, though zero-shot prompting yields the highest aggregate scores across all five models. We further observe systematic differences between models in stylistic alignment with Arabic maqama conventions, and corroborate our findings with a second independent LLM judge, paired statistical significance testing, and non-LLM proxy measures of saj.

Comments14 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑