arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18888cs.CL

评估德语文本自然语言生成中的体验质量

Assessing Quality of Experience in Natural Language Generation of German Text

  • Technische Universität Berlin(柏林工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Dinh Nam Pham, Shushen Manakhimova, Vivien Macketanz, Sebastian Möller

AI总结:

该研究针对德语NLG,构建了含人工评分的TextQ-German数据集,开发出混合等自动QoE预测模型,为NLG评估提供了公开资源与基线,助力贴合人类感知的NLG系统开发。

AI中文摘要:

自然语言生成(NLG)的快速发展使得生成文本的可靠评估愈发关键,因为诸如大语言模型(LLM)这类系统现已广泛部署于实际应用中。然而,传统自动指标无法捕捉感知质量的多面性。本文从体验质量(QoE)视角出发,引入TextQ-German这一新型数据集套件,用于德语NLG的以人为中心评估,涵盖自动文本摘要与机器翻译任务。通过针对德语使用者的众包研究,我们收集了人工质量评分,并为每个任务确定了相关的感知质量维度。我们开发了自动QoE预测模型,包括基于Transformer的模型、基于语言特征的模型及混合方法。在几乎所有实验设置中,混合模型的性能优于纯Transformer基线,而仅使用语言特征即可达到微调语言模型的性能水平。该数据集还补充了带有整体QoE评分标注的LLM生成输出,对保留集的最终验证表明其对未见数据具有泛化能力。本研究提供了可公开获取的NLG评估资源与自动QoE预测基线,为开发更贴合人类质量感知的NLG系统奠定了基础。

英文摘要:

The rapid advancement of Natural Language Generation (NLG) has made the reliable evaluation of generated text increasingly critical, as these systems, such as large language models (LLMs), are now widely deployed in real-world applications. However, traditional automatic metrics fail to capture the multifaceted nature of perceived quality. In this paper, we introduce TextQ-German, a novel dataset suite for human-centered evaluation of German NLG from a Quality of Experience (QoE) perspective, covering automatic text summarization and machine translation. Through crowdsourcing studies with German speakers, we collect human quality ratings and identify relevant perceptual quality dimensions for each task. We develop automatic QoE prediction models, including transformer-based, linguistic feature-based, and hybrid approaches. Hybrid models outperform pure transformer baselines in almost all experimental settings, while linguistic features alone can approach the performance of fine-tuned language models. The dataset is extended with LLM-generated outputs annotated with overall QoE scores. Final validation on held-out sets indicates generalization to unseen data. Our work contributes a publicly accessible resource for NLG evaluation and baselines for automatic QoE prediction, providing a foundation for developing NLG systems that better align with human quality perception.

补充信息

↑