平均偏差:人类忠实性标注并非局部忠实
Averaging Bias: Human Faithfulness Annotations are not Locally Faithful
浏览论文内容
中文总结 AI 辅助
该研究发现文本摘要忠实性基准的人类全局标注存在平均偏差,即标注者接受大部分句子忠实的摘要,而非严格合取规则要求的所有句子都忠实,需改进人类标注设计。
中文摘要 AI 辅助
文本摘要的忠实性评估仅当模型生成摘要的每一句话都得到源文档支持时,才将该摘要判定为忠实,这是一种严格的合取规则,即只要存在一句未被支持的句子,整个摘要就会被判定为不忠实。然而,大多数忠实性基准数据集仅为每个摘要收集一个全局人类标注标签。我们探究这类全局人类标签是否实际遵循该合取规则,假设标注者可能在大部分句子忠实时就接受摘要为忠实,而非要求所有句子都忠实。为验证假设,我们使用五个大语言模型(LLM)作为逐句评估者,在四个广泛使用的忠实性基准数据集上进行实验,发现全局人类标签与逐句LLM判断的平均值的相关性,高于其与严格合取规则实施情况的相关性。人工审查确认,被人类标注为忠实的摘要中,有相当一部分包含真实的局部事实错误,我们将这种倾向称为平均偏差。我们的研究结果表明,广泛使用的忠实性基准上的人类标签存在可测量的平均偏差,呼吁为可靠的人类标注设计精心结构化的方案。
英文摘要
Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjunctive rule under which a single unsupported sentence makes the whole summary unfaithful. Yet most faithfulness benchmarks collect only one global human annotation label per summary. We ask whether such global human labels actually implement the conjunctive rule. We hypothesize that annotators may accept a summary as faithful when most sentences are faithful, not only when all are faithful. To test our hypothesis, we use five large language model (LLM) judges as per-sentence raters across four widely used faithfulness benchmarks. We find that global human labels correlate better with the average of per-sentence LLM judgments than with the implementation of the strict conjunctive rule. A manual review confirms that a substantial fraction of summaries labeled faithful by humans contain genuine local factual errors. We call this tendency Averaging Bias. Our results reveal that human labels on widely used faithfulness benchmarks contain measurable Averaging Bias, calling for carefully structured designs for trustworthy human annotations
发表机构
- Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。