任务分解是否能改进自动自然语言生成(NLG)评估?
Does task decomposition improve automatic NLG evaluation?
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究在多个NLG数据集上对比了有无任务分解的LLMaJ方法,发现任务分解本身不会提升LLMaJ性能,其此前的性能提升源于人工标签训练,无分解的LLMaJ在有人工标签时可与人类标注者表现相当。
AI中文摘要:
LLM-as-a-judge(LLMaJ)框架已成为一种低成本、可复现、无需参考文本的自然语言生成(NLG)评估的有前景解决方案。现有研究尝试通过将评估任务分解为更简单的子任务来改进LLMaJ。本研究在多个NLG数据集上系统比较了采用任务分解与未采用任务分解的LLMaJ方法,发现没有证据表明采用任务分解的LLMaJ相较于公平基线(未使用分解)会带来性能提升;相反,之前报道的基于分解的LLMaJ的性能提升源于使用人工标签作为训练数据,而非任务分解本身。此外,当人工标签可用时,未使用任务分解的LLMaJ可达到与人工标注者相当的性能。
英文摘要:
The LLM-as-a-judge (LLMaJ) framework has emerged as a promising solution for cheap, reproducible, reference-free Natural Language Generation (NLG) evaluation. Prior work seeks to improve LLMaJ by decomposing evaluation tasks into simpler sub-tasks. In this work, we systematically compare LLMaJ methods with and without decomposition on multiple NLG datasets. We find no evidence that LLMaJ with task decomposition leads to performance gains over a fair baseline that does not use decomposition. Instead, we find that previously reported performance gains in decomposition-based LLMaJ stem from using human labels as training data, and not task decomposition itself. Also, we find that, when human labels are available, LLMaJ without using task decomposition can perform comparably to human annotators.