发表机构
Università di Pisa(比萨大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对单人类需在有限审计预算下审计N个LLM智能体的问题,研究了置信度校准不当且存在相关性时的审计预算分配,发现校准阈值的变化规律,验证了部分大语言模型置信度的实用性,给出了无效监督的定量准则。
AI 中文摘要
单个人类需在每轮仅能进行B次审计(B远小于N)的预算下,审计N个LLM智能体,审计时需参考可能被对手故意校准不当的自报告置信度及相关错误。我们将此建模为基于两级高斯 copula 的带预算噪声检测,并确定了校准阈值δ*,超过该阈值后按置信度排序审计的效果比随机审计更差。两个先验预期发生反转:随着预算缩减,δ*上升,且跨家族相关性并非低水平——共同难度优于谱系。五个开源权重LLMs表现出操作上无用(近乎恒定)的置信度,点估计达到或超过翻转点,尽管置信区间跨越该点;一个专有模型具有信息性且低于翻转点。我们给出了“无效”监督的定量准则,在记录的轨迹上回放策略验证了该排序。
英文摘要
A single human must audit $N$ LLM agents under a budget of $B \ll N$ audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the miscalibration threshold $δ^*$ past which confidence-ranked auditing is \emph{worse} than random. Two a-priori expectations reverse: $δ^*$ \emph{rises} as the budget shrinks, and cross-family correlation is not low---shared difficulty dominates lineage. Five open-weight LLMs show operationally useless (near-constant) confidence, point estimates at or beyond the flip though CIs straddle it; a proprietary model is informative and lands below it. We give a quantitative criterion for \emph{vacuous} oversight, and replaying policies on recorded traces confirms the ordering.