arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一致性破坏了共形预测

Conformity Breaks Conformal Prediction

Yibo Hu, Hanyu Su

arXiv 2609.04445首次发表:更新:

发表机构

Illinois Institute of Technology(伊利诺伊理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现多智能体LLM系统中,同伴的错误断言会引发模型评分机制偏移,破坏共形预测,使覆盖率下降,且标准修复方法无效。

AI 中文摘要

当大语言模型(LLM)单独回答问题时,共形保证是有效的;但当同一模型看到其他智能体一致断言错误答案时,共形保证就会失效。问题本身没有改变,只是模型对正确答案的评分发生了变化,我们将这种现象称为评分机制偏移:干净校准能验证模型单独评分时的表现,但无法验证模型在同伴压力下的评分表现。我们表明,这种偏移会悄无声息地破坏多智能体LLM系统中的共形预测。在公开权重模型和多项选择问答任务中,当α=0.10的标准操作点下存在一致错误的同伴时,覆盖率从校准后的90%降至74%。平均数据掩盖了更严重的失败:攻击者通过针对保证仍覆盖的低置信度项目,使该子组的覆盖率几乎减半,从87%降至47%,而监测的平均值仍高得多。这种失败还会影响决策层:本应在不确定时升级处理的系统,反而可能变得足够自信,根据攻击者的错误答案采取行动。标准的共形修复方法无法解决该问题,因为问题分布没有改变,改变的是模型的评分行为。

英文摘要

A conformal certificate can be valid when an LLM answers alone and invalid when the same LLM sees peers that unanimously assert a wrong answer. The question is unchanged; the model's score for the correct answer changes. We call this a score-mechanism shift: clean calibration certifies how the model scores answers alone, but not how it scores them under peer pressure. We show that this shift silently breaks conformal prediction in multi-agent LLM systems. Across open-weight models and multiple-choice QA tasks, coverage falls from a calibrated 90% to 74% under unanimous-wrong peers at the standard alpha = 0.10 operating point. The average hides a sharper failure: by targeting the low-confidence items the certificate still covers, an attacker nearly halves coverage on that subgroup, from 87% to 47%, while the monitored average remains much higher. The failure also reaches the decision layer: a system that should escalate when uncertain can instead become confident enough to act on the attacker's wrong answer. Standard conformal fixes do not solve the problem, because the question distribution has not changed; the model's scoring behavior has.

Comments19 pages, 6 figures, 11 tables. Code: https://github.com/yibo-hu-lab/conformity-breaks-conformal

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑