伪造的同行判断会误导多模态大语言模型(LLM)评判小组:无来源锚定与小组共识验证
Forged Peer Judgments Mislead Multimodal LLM Judge Panels: Source-Blind Anchoring and Panel-Consensus Verification
查看机构详情
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究发现多模态LLM评判小组存在无来源锚定的文本攻击面,伪造同行判断可误导判决,提出的小组共识验证方法能有效阻断此类攻击并降低损害。
中文摘要 AI 辅助
多模态大语言模型(LLM)评判小组可交叉参考同行意见,但引用的同行判断本身可能不可靠。我们在视觉语言模型(VLM)评判小组中发现了无来源锚定这一文本层面的攻击面:在自我评判和同行评判框架下,引用独立的视觉判断会产生19至26个百分点的巨大锚定差距。内容匹配、仅标签不同的对照组的错误率仅变化了-0.17个百分点(95%置信区间[-0.68,0.35]),表明自我/同行标签本身无法解释该效应。在我们测试的构造下,刻意生成的简洁错误引用推翻原本正确判决的频率是自然出现的错误同行陈述的1.5至2.7倍,在两个数据集和七个VLM评判器上,自助法95%置信区间不包含奇偶性。由于两种陈述群体在选择和形式上存在差异,该比率衡量的是测试攻击下的差异损害,而非仅来源的因果效应。随后我们提出了小组共识验证方法,即对照独立收集的盲投票核查引用内容,该方法可阻断84.9%的伪造攻击,将其净损害降低97.5%,并在留一法重新验证下保留了真实同行信息的积极但统计上无定论的点估计。这些结果确定了一个低成本攻击面和一种用于更安全的多模态协作评估的具体防御方法。
英文摘要
Multimodal LLM judge panels can cross-reference peers, but a quoted peer judgment may itself be untrusted. We expose source-blind anchoring as a text-level attack surface in vision-language model (VLM) panels. Quoting independent visual judgments creates large anchoring gaps (19--26 percentage points) under both self and peer framing. A matched-content, label-only control changes the broken rate by only $-0.17$pp (95\% CI $[-0.68,0.35]$), showing that the self/peer label itself does not explain the effect. Under our tested construction, deliberately generated, concise wrong quotes overturn originally-correct verdicts 1.5--2.7$\times$ more often than naturally occurring wrong peer statements, with bootstrap 95\% CIs excluding parity across two datasets and seven VLM judges. Because the two statement populations differ in selection and form, this ratio measures differential damage under the tested attack rather than a provenance-only causal effect. We then introduce panel-consensus verification, which cross-checks a quote against independently collected blind votes. It blocks 84.9\% of fabricated attacks, cuts their net harm by 97.5\%, and preserves the positive but statistically inconclusive point estimate for genuine peer information under leave-one-out re-verification. These results identify a low-cost attack surface and a concrete defense for safer multimodal collaborative evaluation.