arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用LLM评估可读性:推理与少样本提示的作用

Assessing Readability with LLMs: The Role of Reasoning and Few-Shot Prompting

Raphaël Thieffry, Matej Martinc

arXiv 2609.24650首次发表:更新:

发表机构

Université Paris-Saclay; Jožef Stefan Institute(巴黎-萨克雷大学; 约热夫·斯特凡研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统评估了多种开源LLM在多语言可读性评估中的表现,发现思维链推理和少样本提示显著提升预测质量,验证了LLM作为跨语言可读性评估器的有效性。

AI 中文摘要

可读性评估对于在教育、医疗保健和信息检索等领域为预期受众定制文本至关重要。然而,传统的可读性公式难以在不同体裁和语言间泛化,而监督式机器学习模型依赖于稀缺的、特定领域的标注语料库,限制了其适用性——尤其是对于资源较少的语言。大型语言模型(LLMs)提供了一种高度可扩展、多语言的替代方案,无需任务特定训练,但高级提示策略对其性能的影响仍未得到充分探索。在本文中,我们对多种开源LLMs进行了系统基准测试,用于多语言可读性评估,重点关注教育框架所需的离散可读性级别的预测。除英语外,我们还在一种资源较少的语言——斯洛文尼亚语上评估了我们的方法,以确定LLMs在低资源环境下是否仍然有效。具体而言,我们研究了显式推理的影响,证明思维链(CoT)提示和面向推理的模型相比直接回答能带来显著改进。此外,我们对少样本上下文学习的探索表明,与零样本设置相比,每类别仅提供一个标注示例(1-shot)即可大幅提升预测质量,而额外示例带来的收益递减。通过将这些方法与传统无监督指标和最先进的监督基线进行全面比较,我们确立了开箱即用的LLMs作为稳健、跨语言可读性评估器的可行性。

英文摘要

Readability assessment is essential for tailoring texts to intended audiences across educational, healthcare, and information retrieval domains. However, traditional readability formulas struggle to generalize across genres and languages, while supervised machine learning models rely on scarce, domain-specific annotated corpora, limiting their applicability--particularly for less-resourced languages. Large Language Models (LLMs) offer a highly scalable, multilingual alternative that requires no task-specific training, yet the impact of advanced prompting strategies on their performance remains underexplored. In this paper, we conduct a systematic benchmark of diverse open-source LLMs for multilingual readability assessment, focusing on the prediction of discrete readability levels required by educational frameworks. In addition to English, we evaluate our approach on a less-resourced language, Slovenian, to establish whether LLMs remain effective in low-resource settings. Specifically, we investigate the influence of explicit reasoning, demonstrating that Chain-of-Thought (CoT) prompting and reasoning-oriented models yield significant improvements over direct answering. Furthermore, our exploration of few-shot in-context learning reveals that providing just one labelled example per category (1-shot) substantially enhances prediction quality compared to zero-shot settings, with additional examples offering diminishing returns. By comprehensively comparing these approaches against traditional unsupervised metrics and state-of-the-art supervised baselines, we establish the viability of out-of-the-box LLMs as robust, cross-lingual readability assessors.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑