arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人类与大语言模型(LLM)的标注挑战:评价性语言的案例研究

Challenges in annotations by humans and LLMs: A case study of evaluative language

Mirela Imamovic, Aenne Cecilia Kristine Knierim, Khushi Pitroda, Ekaterina Lapshinova-Koltunski

arXiv 2607.28119首次发表:更新:

AI 中文总结

本文以TED演讲语料库为对象,对比受训语言学者、专业学者与LLM在评价性语言标注上的表现,发现LLM经微调后F1值达0.77,表现优于专业学者,可辅助复杂标注任务。

AI 中文摘要

本文对比了受训语言学者、专业语言学者以及大语言模型(LLM)生成的标注,探究它们是否在复杂语言现象处理上存在相似困难。为此,我们分析了英语口语化科普语篇中的评价性语言,以英文TED演讲文稿语料库为例,聚焦评价理论及其态度子系统,涵盖情感(Affect)、判断(Judgement)、鉴赏(Appreciation)三类范畴。评价理论属于高度主观的标注任务,适合研究复杂标注挑战。首先,我们评估特定科学领域句子层面的人类标注;接着,设计三个提示词,对比模型在评价类自动分类任务中的表现;使用最优提示词评估三个LLM的性能并对模型进行微调,最终达到0.77的F1值。研究发现,与专业语言学者的标注相比,模型表现最佳,而受训语言学者未达到高一致性分数。我们得出结论:LLM可辅助解决复杂标注任务,为数字人文研究中复杂理论的标注与分析开辟新路径。

英文摘要

In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑