发表机构
Illinois Institute of Technology; Amazon(伊利诺伊理工学院; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现,在LLM安全评审团中,同行的错误标签断言会大幅提升误报率,且评审员更易追随“不安全”方向,多数投票机制会因社交压力失效,还提供了部署前诊断方法。
AI 中文摘要
大语言模型(LLMs)越来越多地被用于检测不安全内容,一种常见方法是结合评审团中多个模型的判断来修正单个模型的错误,但当每个模型在投票前都看到相同的误导性上下文时,这种优势可能会消失。我们在受控的两轮实验中研究了这一风险:每个模型先单独判断一个项目,然后在6个模拟同行要么断言错误标签、要么弃权(不执行)后再次判断该项目,我们通过多数投票结合最终判断。在6个开放权重LLMs和6个数据集上,我们发现断言错误标签的同行消息会将平均评审员误报率从同行沉默时的56.5%提升至87.5%,且多数投票会将评审团的误报率提升至100%;若无断言标签,同一评审团的表现优于其平均成员。该效应具有强烈不对称性:评审员追随“不安全”方向推动的比例(约75%)远高于“安全”方向(约17%),因此评审团的误报率急剧上升,而漏报率变化很小。专有模型探测显示不同模型间存在显著差异。这些结果表明,对共享社交线索的易感性是安全评审团的一种失效模式,并提供了一种简单的部署前诊断方法。
英文摘要
Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but this benefit may disappear when every model sees the same misleading context before voting. We study this risk in a controlled two-round experiment. Each model first judges an item alone, then judges it again after six simulated peers either assert the wrong label or abstain. We combine the final judgments by majority vote. Across six open-weight LLMs and six datasets, we find that the wrong-label peer message raises the average reviewer false-alarm rate from 56.5% under silent peers to 87.5%, and majority voting raises the panel false-alarm rate to 100%. Without an asserted label, the same panel outperforms its average member. The effect is strongly asymmetric: reviewers follow pushes toward "unsafe" far more than pushes toward "safe" (about 75% versus 17%), so the panel's false-alarm rate rises sharply while its harmful-miss rate changes little. The proprietary-model probe shows substantial variation across models. These results identify susceptibility to shared social cues as a failure mode of safety panels and provide a simple pre-deployment diagnostic.