骨干网络演进对基于LLM的相关性评估的影响
The Impact of Backbone Evolution on LLM-Based Relevance Assessments
浏览论文内容
中文总结 AI 辅助
本研究通过固定提示词评估不同版本LLM骨干网络的相关性评判性能,发现新版本不必然提升评判质量,且早期正确判断可能不被保留,警示提示词跨版本迁移的假设。
中文摘要 AI 辅助
大语言模型(LLM)正在快速发展,新模型展现出更强的能力。这表明,在基于LLM的相关性评判中,在相同提示词下,能力更强的模型将实现与人类判断更高的一致性。我们通过研究基于LLM的相关性评判器在骨干网络演进下的行为来挑战这一理解。在保持提示词不变的情况下,我们评估了一个代表性的单提示词(UMBRELA)和一个基于评分标准的提示词(EXAM),跨越商业(Gemini、GPT)和开放权重(Qwen、Llama)模型的连续模型版本。总体而言,我们没有发现一致的证据表明新版本能带来更好的相关性评判器。至关重要的是,相似或改进的总体性能并不意味着判断的稳定性:早期版本的LLM骨干网络做出的正确判断不一定被后续版本保留。我们调查了这些退化的潜在驱动因素。我们的发现警示人们不要假设为某一骨干网络版本设计和验证的评判提示词在模型更新后会表现相当或更好,即使在同一模型家族内也是如此。
英文摘要
LLMs are evolving rapidly, with newer models offering stronger capabilities. This suggests that in LLM-based relevance judging, more capable models will achieve higher agreement with human judgements under the same prompt. We challenge this understanding by investigating the behavior of LLM-based relevance judges under backbone evolution. Keeping the prompts fixed, we evaluate a representative single-prompt (UMBRELA) and a rubric-based prompt (EXAM) across sequential model versions of commercial (Gemini, GPT) and open-weight (Qwen, Llama) models. Overall, we find no consistent evidence that newer versions lead to better relevance judges. Crucially, similar or improved aggregate performance does not imply judgment stability: correct judgements made by an earlier version of an LLM backbone are not necessarily preserved by later versions. We investigate the potential drivers of these regressions. Our findings caution against the assumption that judging prompts designed and validated for one backbone version will perform equivalently or better when the model is updated, even within the same family.
发表机构
- The University of Queensland(昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。