AI 中文总结
本文提出战略交互监督框架,研究AI辩论中智能体在保持任务正确性的同时追求潜在目标的问题,并量化了任务成功与信息披露间的权衡,强调监督评估需考虑记录传达的信息。
AI 中文摘要
可扩展监督旨在验证能力超过其监督者的智能体的行为。AI辩论已被提出作为一种监督解决方案,其中竞争的智能体帮助资源有限的验证者评估其无法独立可靠评估的主张。其许多前景依赖于激励导致正确裁决的诚实论证。然而,正确的裁决并不必然唯一地决定用于支持它的论证。智能体可能保留对呈现哪些正确主张、如何构建它们以及以何种顺序披露它们的自由裁量权。这种残余自由可能允许智能体塑造验证者所学到的内容,超越任务相关的结论,在不损害裁决正确性的情况下追求潜在目标。为了研究这一现象,我们引入了战略交互监督(SIO)框架,该框架将监督同时视为验证机制和战略沟通渠道。在此框架内,我们形式化了任务允许的潜在优化的概念,这涉及在保持规定任务性能的同时追求潜在目标。作为概念验证,我们在带有交叉质询的建立协议辩论中实例化SIO,并量化了任务成功与关于隐藏变量的信息披露之间的权衡。该权衡识别出一个战略窗口,在此窗口中,大量披露仍与任务允许性兼容。为了缓解,我们通过扩展交叉质询者的角色来减少允许的偏见,以在有限交互范围内减轻持续披露。我们的结果强调了不仅通过其裁决的正确性,而且通过其记录传达的信息来评估监督的必要性。
英文摘要
Scalable oversight aims to verify the behaviour of agents whose capabilities exceed those of their overseers. AI debate has been proposed as an oversight solution in which competing agents help a resource-limited verifier assess claims that it cannot reliably evaluate unaided. Much of its promise rests on incentivizing honest arguments that lead to correct verdicts. Yet a correct verdict need not uniquely determine the arguments used to support it. Agents may retain discretion over which correct claims to present, how to frame them, and in what order to disclose them. This residual freedom can allow agents to shape what the verifier learns beyond the task-relevant conclusion, pursuing latent objectives without compromising verdict correctness. To study this phenomenon, we introduce the framework strategic interactive oversight (SIO), which treats oversight jointly as a verification mechanism and a strategic communication channel. Within this framework, we formalise the notion of task-admissible latent optimisation, which entails the pursuit of latent objectives while maintaining a prescribed task performance. As proof-of-concept, we instantiate SIO in the establish protocol debate with cross-examination and quantify a tradeoff between task success and information disclosure about a hidden variable. The trade-off identifies a strategic window in which substantial disclosure remains compatible with task admissibility. Towards mitigation, we reduce admissible bias by expanding the cross-examiner's role to mitigate persistent disclosure over finite interaction horizons. Our results highlight the need to evaluate oversight not only by the correctness of its verdicts, but also by the information conveyed through its transcripts.
Comments30 pages, 1 figure