置信路由实际上在做什么:多智能体 deliberation 中的路由、校准与承诺审计
What Confidence Routing Is Actually Doing: Auditing Routing, Calibration, and Commitment in Multi-Agent Deliberation
- Argonne National Laboratory(阿贡国家实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文审计多智能体置信路由协议,分离路由、校准与承诺三个维度,发现置信度虽能区分候选(AUROC 0.72)但过度自信,校准可改善而区分能力缺失,且承诺偏差因设置而异,强调部署前需分别评估三者。
AI中文摘要:
一种常见的多智能体设计是让智能体报告置信度,并让得分最高的智能体接着发言,这隐式地使用一个标量同时用于路由对话和估计不确定性。我们通过分离三个轨迹层面的问题来审计这种置信路由广播协议:它是否选择了正确的候选者(路由),报告的置信度是否表现得像概率(校准),以及被选中的智能体是否公开陈述了赢得回合的答案(承诺)。我们的主要研究覆盖了 4,181 条 gpt-oss-120b olympiad-math 轨迹;我们在一个 2×2 的智能体×基准网格上重复了审计,该网格增加了 gemma-4-31B-it 和一个生物学多项选择基准。在主单元格中,置信度能区分正确与错误的候选者(AUROC 0.72),但严重过度自信(平均陈述置信度 79% 对比准确率 52%)。一种交叉拟合、分层分层的等渗回归程序将留出候选者的期望校准误差从 0.278 降至 0.008,但它无法恢复缺失的区分能力:在两个 Gemma 单元格中,原始 AUROC 仅为 0.537 和 0.440。路由同样依赖于设置。在 gpt-oss/math 上,固定路由器的差异最多为 1.1 个百分点,而在 Gemma 单元格中,原始置信度 argmax 比随机有效选择低 5.6 和 11.2 个百分点。承诺又有所不同:在主单元格中,投票和口头答案在 20.4% 的有效对中不一致,其中 62.4% 的修订是新生成的,无条件正确性变化为 -1.7 个百分点;其他三个单元格则从 +0.9 到 +12.2 个百分点不等。可迁移的经验是程序性的:在原始置信度用于部署决策之前,必须分别测量路由区分能力、概率校准和公开承诺。
英文摘要:
A common multi-agent design asks agents to report confidence and lets the highest-scoring agent speak next, implicitly using one scalar both to route the conversation and to estimate uncertainty. We audit this confidence-routed broadcast protocol by separating three trace-level questions: whether it selects the right candidate (routing), whether reported confidence behaves like a probability (calibration), and whether the selected agent publicly states the answer that won the turn (commitment). Our primary study covers 4,181 gpt-oss-120b olympiad-math traces; we repeat the audit on a 2-by-2 actor-by-benchmark grid that adds gemma-4-31B-it and a biology multiple-choice benchmark. In the primary cell, confidence discriminates correct from wrong candidates (AUROC 0.72) but is strongly overconfident (79% mean stated confidence versus 52% accuracy). A cross-fitted, tier-stratified isotonic procedure reduces Expected Calibration Error from 0.278 to 0.008 on held-out candidates, but it does not recover missing discrimination: raw AUROC is only 0.537 and 0.440 in the two Gemma cells. Routing is likewise setting-dependent. Fixed routers differ by at most 1.1 percentage points on gpt-oss/math, whereas raw-confidence argmax performs 5.6 and 11.2 points below random-valid selection in the Gemma cells. Commitment is distinct again: in the primary cell, poll and spoken answers diverge in 20.4% of valid pairs, 62.4% of those revisions are fresh generations, and the unconditional correctness shift is -1.7 points; the other three cells instead range from +0.9 to +12.2 points. The transferable lesson is procedural: routing discrimination, probability calibration, and public commitment must be measured separately before raw confidence is used for deployment decisions.