arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28636cs.CLcs.CY

模型链:面向偏见鲁棒性大语言模型评判器的跨模型审计

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

Qian Wang, Zhanzhi Lou, Zhenheng Tang, Nuo Chen, Bingsheng He

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对LLM评判器的认知偏见问题,提出跨模型审计流程CoM,发现审计器身份的关键作用,制定偏见特异性审计器选择规则,在多模型多偏见实验中取得优于基线的准确率。

中文摘要 AI 辅助

大语言模型(LLM)日益成为自动化评判器,但其判断仍易受认知偏见影响。现有缓解措施大多依赖提示驱动的去偏见方法,该方法对不同偏见类型鲁棒性差,或依赖人工评估,无法规模化。我们研究了模型链(Chain-of-Models, CoM),这是一种自动化审计流程,其中第二个模型在生成最终判断前检查第一个模型的推理轨迹。核心设计问题是审计器应使用相同模型、同系列模型还是不同系列模型。在6个系列的9个模型、4种认知偏见和4个事实数据集上,我们发现审计器身份在两方面起关键作用:其一,独立抗偏见能力无法预测审计有效性:Kimi-K2.5在多种偏见中是最强的独立模型,但作为Qwen2.5-72B偏见轨迹的审计器时表现较弱;其二,最佳审计器具有偏见特异性:GPT-4o在从众、权威和分心偏见中表现最强,而GLM-5在谄媚偏见中表现最强。我们基于这些发现,制定了针对每种偏见的审计器选择规则,给定偏见类型后,该规则会依据功能多样性、每种偏见的独立抗偏见能力以及校准后的审计有效性对候选者打分。在校准/测试划分下,该选择器在四个偏见切片上达到最高准确率(0.884,相比最强单一固定审计器的0.824和无审计基线的0.805)。我们在该httpsURL发布了数据、配置和一项LLM智能体技能。

英文摘要

LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittle across bias types, or human evaluation, which does not scale. We study \emph{Chain-of-Models} (CoM), an automated audit pipeline in which a second model inspects the first model's reasoning trace before producing the final judgment. The key design question is whether the auditor should be the same model, a same-family model, or a different-family model. Across 9 models from 6 families, 4 cognitive biases, and 4 factual datasets, we find that auditor identity matters in two ways. First, standalone bias resistance does not predict audit effectiveness: Kimi-K2.5 is the strongest standalone model on several biases, yet is a weak auditor for Qwen2.5-72B's biased traces. Second, the best auditor is bias-specific: GPT-4o is strongest on bandwagon, authority, and distraction, while GLM-5 is strongest on sycophancy. We operationalize these findings with a per-bias auditor selection rule that, given the bias type, scores candidates along functional diversity, per-bias standalone resistance, and calibrated audit effectiveness. Under a calibration/test split, the selector reaches the highest accuracy across the four biased slices ($0.884$ vs.\ $0.824$ for the strongest single fixed auditor and $0.805$ for the no-audit baseline). We release data, configurations, and an LLM-agent skill at https://anonymous.4open.science/r/chain-of-models-B585 .

补充信息

↑