arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于提升大语言模型加泰罗尼亚语文本简化能力的强化学习

Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities

Arnau Ayguadé Domingo, Stefan Bott, Horacio Saggion

arXiv 2609.04823首次发表:更新:

发表机构

Universitat Pompeu Fabra; Barcelona Supercomputing Center(庞培法布拉大学; 巴塞罗那超级计算中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对低资源语言加泰罗尼亚语,提出结合SARI指标与惩罚项的GRPO奖励函数,后训练IberianLLM-7B-Instruct提升其ATS性能,同时探索跨语言迁移学习但未获显著效果。

AI 中文摘要

尽管自动文本简化(ATS)对无障碍访问至关重要,但其进展并未跟上更广泛自然语言处理技术的快速发展。本文研究将强化学习(RL)应用于大语言模型(LLMs),以提升低资源语言的ATS质量。本文提出一种新颖的奖励函数,用于引导LLMs采用目标简化风格,结合SARI指标与特定惩罚项,采用组相对策略优化(GRPO)实现。通过在ASSET数据集上对IberianLLM-7B-Instruct进行后训练,验证了该GRPO及奖励函数的有效性。在英文ASSET上后训练后,该模型在两个精选加泰罗尼亚语基准上的ATS性能提升,同时成功抑制了此前观察到的负面行为。本文还探索了跨语言迁移学习,将ASSET分别翻译成加泰罗尼亚语和西班牙语并对模型进行后训练,但这些操作未在域外交代基准上显示出显著提升。

英文摘要

Although automatic text simplification (ATS) is critical for accessibility, its progress has not matched the rapid evolution of broader natural language processing techniques. This paper investigates the application of reinforcement learning (RL) to improve the quality of ATS for low-resource languages using Large Language Models (LLMs). The paper introduces a novel reward function, designed to guide LLMs toward a targeted simplification style with Group Relative Policy Optimization (GRPO), that combines the SARI metric with specific penalty components. The effectiveness of GRPO with this reward function is motivated and demonstrated by post-training IberianLLM-7B-Instruct on the ASSET dataset. After post-training on the English ASSET, the model's ATS performance improves on two curated Catalan benchmarks while also successfully suppressing previously observed negative behaviors. Cross-lingual transfer learning is explored by translating ASSET into Catalan and Spanish and post-training the model on each version, but these fail to show a significant improvement on the out-of-domain benchmark.

CommentsAccepted at CLEAR-TEXT 2026: Readability and text simplification workshop at the International Conference Computational Linguistics in Bulgaria (CLIB 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑