多智能体辩论蒸馏中认知可靠性的黑盒审计
Black-Box Auditing of Epistemic Reliability in Multi-Agent Debate Distillation
AI总结:
针对多智能体辩论蒸馏中的认知可靠性退化,提出黑盒审计框架ER-Audit,通过两阶段反例搜索与序贯检验揭示被性能增益掩盖的隐藏任务退化。
AI中文摘要:
辩论蒸馏利用多智能体辩论记录来调整较弱的验证器,以提高其在后续辩论中的判断能力,但在受监控任务上的增益并不能确立其在相关未受监控任务上的可靠性。我们研究认知可靠性退化问题,即调整在保持受监控性能的同时,降低了对隐藏任务上正确响应的支持。我们考虑一个对抗性辩论者,它在捍卫正确的受监控响应的同时操纵辩论论点,并询问由此产生的退化是否仅仅反映了灾难性遗忘,以及标准评估能否检测到它。为了解决这些问题,我们提出了ER-Audit,一个两阶段的黑盒审计框架,用于比较调整前后的冻结验证器检查点,并引入了两个评估基准,将基于共享上下文的受监控和隐藏任务提示配对。ER-Audit通过评估语义有效的改写来搜索非退化的反例,如果找不到,则使用独立的改写进行序贯假设检验。我们推导出非退化概率的随时有效下置信界,允许在有限预算内进行数据依赖的停止。我们进一步在固定改写分布上建立了一个公共下界,并将其扩展到与其混合物的总变差距离有界的分布。我们的实验表明,更高的隐藏任务准确率可以与更多的非退化反例和更低的非退化下界共存。这种分歧挑战了仅基于广泛灾难性遗忘的解释,并表明审计可以揭示被聚合性能增益掩盖的选择性隐藏任务退化。我们的代码和基准可在以下网址获取:https://this https URL。
英文摘要:
Debate distillation adapts weaker verifiers using multi-agent debate transcripts to improve their judgement in subsequent debates, but gains on monitored tasks do not establish reliability on related unmonitored tasks. We study epistemic reliability degradation, in which adaptation preserves monitored performance while reducing support for correct responses on hidden tasks. We consider an adversarial debater that manipulates debate arguments while defending the correct monitored response, and ask whether the resulting degradation merely reflects catastrophic forgetting and whether standard evaluation can detect it. To address these questions, we propose ER-Audit, a two-stage black-box auditing framework that compares frozen verifier checkpoints before and after adaptation, and introduce two evaluation benchmarks pairing monitored and hidden task prompts grounded in shared contexts. ER-Audit searches for counterexamples to non-degradation by evaluating semantically valid paraphrases and, if none is found, uses independent paraphrases for sequential hypothesis testing. We derive anytime-valid lower confidence bounds on the non-degradation probability, allowing data-dependent stopping within a finite budget. We further establish a common lower bound across fixed paraphrase distributions and extend it to distributions within a bounded total variation distance of their mixtures. Our experiments show that higher hidden-task accuracy can coexist with more counterexamples to non-degradation and lower non-degradation bounds. This divergence challenges explanations based solely on broad catastrophic forgetting and shows that auditing can uncover selective hidden-task degradation concealed by aggregate performance gains. Our code and benchmarks are available at https://github.com/CSIRO-CQS-AI-alignment-Team/Epistemic-Reliability-Auditor.