发表机构
National Institute of Informatics (NII); University of Amsterdam; University of Tokyo; Vrije Universiteit Amsterdam; Centrum Wiskunde & Informatica (CWI); Utrecht University(信息学研究所(NII); 阿姆斯特丹大学; 东京大学; 阿姆斯特丹自由大学; 数学和计算机科学中心(CWI); 乌得勒支大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大模型作为评判者系统,发现其信任评分与真相判决的分离性弱于人类,来源线索会同时影响信任评分与真相判决,因此不应将信任评分视为真相判断的独立证据。
AI 中文摘要
大模型作为评判者(LLM-as-Judge)系统可生成多维度评估,如可信赖性、可靠性和事实性,这些输出常被解读为独立证据。针对信任评分与二元真相分类这一常见判断对,我们在正确性受控的问答(QA)任务中检验该假设,发现大模型评判者的信任评分与真相判决的一致性强于人类行为参考,表明信任与真相判断的分离性更弱。我们通过仅改变人类与人工智能相同问答的来源线索开展压力测试,来源归属不仅改变信任评分,还改变真相判决与对数几率衍生的正确侧概率。结果显示,当前大模型作为评判者的协议不应将信任评分视为真相判断的独立证据。
英文摘要
LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidence. We test this assumption for a common pair of judgments: trust scoring and binary truth classification. On correctness-controlled QA, LLM judges align trust scores with truth verdicts more tightly than human behavioral reference, suggesting weaker separations between trust and truth judgment. We then apply stress tests by changing only source cues of identical QA between Human and AI. Source attribution shifts not only trust scores but also truth verdicts and logit-derived correct-side probabilities. Results show that current LLM-as-Judge protocols should not treat trust scores as independent evidence for truth judgments.