前沿大语言模型中的响应漂移
Response drift across frontier large language models
浏览论文内容
中文总结 AI 辅助
研究前沿大语言模型的响应漂移问题,通过47名参与者对十个模型的62个多领域问题进行盲测评估,发现各模型漂移程度不同,依赖领域和问题,且自动指标解释方差低,揭示需以人类评估了解漂移情况。
中文摘要 AI 辅助
所有前沿大语言模型(LLMs)都存在响应漂移——产生与专家验证参考不一致的输出——但这种漂移的程度和结构仍未通过系统的人工评估来表征。在此我们报告了一项完全交叉评估,47名来自不同地理位置的参与者在盲测条件下对十个前沿LLMs的所有62个多领域问题进行评估,得到29140个独立评估。每个模型都有漂移,幅度差异很大:八个模型收敛到统计学上无法区分的上限(偏差78 - 81%),两个偏差较低(47 - 49%)。六个领域和62个问题的漂移情况不同,上限模型之间的成对相关性超过r = 0.85。自动相似性指标在人类判断中的方差解释率不到2%。这些发现表明响应漂移在前沿LLMs中普遍存在,结构上依赖领域和问题,且只能通过以人类为中心的评估来了解。
英文摘要
All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet the magnitude and structure of this drift remain uncharacterised by systematic human evaluation. Here we report a fully crossed evaluation in which 47 geographically diverse participants each assessed all 62 multidomain questions across ten frontier LLMs under blinded conditions, yielding 29,140 independent assessments. Every model drifts, but drift magnitude varies substantially: eight models converge on a statistically indistinguishable ceiling (78-81% deviation), while two achieve lower deviation (47-49%). Drift profiles differ across six domains and 62 questions, with pairwise correlations among ceiling models exceeding r = 0.85. Automated similarity metrics explain less than 2% of variance in human judgements. These findings reveal that response drift is universal across frontier LLMs, domain- and question-dependent in structure, and accessible only through human-centred evaluation.
发表机构
- University of North Texas(北德克萨斯大学)
- Fordham University(福特汉姆大学)
机构由 AI 辅助整理,请以论文原文为准。