发表机构
The University of Queensland(昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以角色设定为探测工具,分析基于大语言模型的IR评估中评估者敏感性,发现高容量模型能保持系统排序一致性,角色来源影响小于评估者角色和模型容量。
AI 中文摘要
大语言模型(LLMs)正越来越多地被用作信息检索(IR)评估中的相关性评估者,这引发了关于评估者框架如何影响判断可靠性及下游系统比较的问题。本研究将角色设定(persona conditioning)作为一种诊断机制,用于揭示LLM评估者的敏感性。我们使用来自两个互补来源(PersonaHub和NVIDIA Nemotron-Personas-USA)的面向任务的角色设定,实例化了五个评估者角色,分别强调意图解读、领域专业知识、对比判断、证据验证以及全局搜索质量评估,并将其与标准的UMBRELA基线进行对比。在TREC DL20和RAG24数据集上,针对六个LLM骨干模型开展的分析显示,评估者敏感性呈现结构化而非均匀的特征。判断结果通常与基线接近,但会在评估严格度、证据阈值或解读重点上发生变化,而非产生大范围的相关性反转。在系统层面,高容量模型能保持系统排序的一致性,而较小规模的模型会放大角色设定引发的不稳定性。局部排名位移分析表明,敏感性集中在特定的检索系统及系统类型上,尤其是DL20上的神经排序/重排序系统和RAG24上的RAG导向管道。角色来源的影响小于评估者角色和模型容量。这些发现表明,角色设定式判断可作为一种受控的敏感性探测工具,用于对基于LLM的IR评估管道进行压力测试,并识别出评估结果对评估者框架敏感的系统。
英文摘要
Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We study persona conditioning as a diagnostic mechanism for exposing LLM assessor sensitivity. Using task-oriented personas drawn from two complementary sources (PersonaHub and NVIDIA Nemotron-Personas-USA), we instantiate five assessor roles emphasizing intent interpretation, domain expertise, contrastive judgment, evidence verification, and global search-quality assessment, compared with a standard UMBRELA baseline. Across six LLM backbones on TREC DL20 and RAG24, our analyses reveal structured rather than uniform assessor sensitivity. Judgments usually remain close to the baseline while shifting assessment strictness, evidential threshold, or interpretation emphasis rather than producing widespread relevance reversals. At the system level, high-capacity models preserve system-ranking agreement, while smaller models amplify persona-induced instability. Local rank-displacement analysis shows sensitivity concentrates on particular retrieval systems and system types, especially neural ranking/reranking systems on DL20 and RAG-oriented pipelines on RAG24. Persona source matters less than assessor role and model capacity. These findings position persona-conditioned judging as a controlled sensitivity probe for stress-testing LLM-based IR evaluation pipelines and identifying systems whose evaluation outcomes are sensitive to assessor framing.
CommentsAccepted at CIKM 2026