发表机构
University of Maryland; Johns Hopkins University(马里兰大学; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文在科学问答场景下,通过混合效应模型研究引用对人类与四个开源LLM偏好的影响,发现人类偏好多样且数量少的引用,LLM也有相关偏好,还探讨了对偏好数据收集的启示。
AI 中文摘要
许多自然语言处理(NLP)任务要求系统在输出中提供归因信息,即对基础来源的引用。归因可作为抵御模型幻觉的壁垒,也可供用户验证模型输出的可信度。然而,当人类和大型语言模型(LLM)比较输出时如何评价引用,这一对于奖励建模和现代LLM后训练至关重要的过程,目前尚不明确。本文在科学问答语境下,研究引用在人类评判者与四个开源LLM偏好中的作用,利用混合效应模型探究引用对成对判断的影响。核心发现包括:(1)人类偏好更多样化但数量更少的引用;(2)尽管无法访问引用来源,LLM相比人类仍表现出一些与引用相关的偏好,但这些偏好依赖于数据和特定模型。我们进一步探讨了研究结果对偏好数据收集的启示。
英文摘要
Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding sources. Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs. Yet, it is unclear how humans and LLMs evaluate citations when comparing outputs, a process central to reward modeling and modern LLM post-training. This paper studies the role of citations in the preferences of human judges and four open-source LLMs within the context of scientific question answering, leveraging mixed effects models to investigate the influence of citations on pairwise judgments. Among our key findings are (1) that humans prefer more diverse citations but fewer overall, and (2) that LLMs show some citation-related preferences compared to humans, despite lacking access to the sources, but these preferences depend on the data and specific models. We further discuss the implications of our findings for preference data collection.