如何进行一场敏感的辩论:一种针对AI辩论的实例最优协议
How to Have a Sensitive Debate: An Instance-Optimal Protocol for AI Debate
浏览论文内容
中文总结 AI 辅助
本文提出一种针对AI辩论的实例最优协议,通过稳定分解与分数块敏感性关联,实现最坏情况正确性、主导策略均衡及黑盒下界最优性。
中文摘要 AI 辅助
随着强大的人工智能系统在一系列认知要求较高的任务中达到甚至有时超越人类专家的能力,对这些系统进行准确监督和管理的问题变得越来越紧迫。一种有前景的方法是AI辩论,旨在利用两个强大AI之间的辩论,将复杂问题分解为更容易直接判断的简单声明。关于辩论的理论工作已用计算复杂性理论的语言形式化了这一直觉,其目标是设计协议(即辩论游戏的规则),为在有限监督下判断复杂问题的解决方案提供严格的正确性保证。具体来说,当前最佳协议已被证明适用于所有具有足够稳定的子问题分解的问题。在本文中,我们为同一类问题设计了一种新协议,在几个方面改进了先前的工作。首先,正确性在最坏情况而非平均情况下成立。其次,诚实和正确是双方辩论者的主导策略均衡,而非斯塔克尔伯格均衡。最后,我们证明了黑盒下界,表明我们的新协议在实例上是最优的。也就是说,对于这类问题,任何仅对人类判断进行黑盒查询的协议都无法超越我们的协议。我们通过将稳定问题分解的概念与查询复杂性中的分数块敏感性概念联系起来,获得了这些结果。
英文摘要
As powerful AI systems reach and sometimes surpass the abilities of human experts across a range of cognitively demanding tasks, the problem of accurate oversight and supervision of these systems has become increasingly urgent. One promising approach is AI debate, which seeks to leverage a debate between two powerful AIs to break complex questions down into simpler claims that can be easily judged directly. Theoretical work on debate has formalized this intuition in the language of computational complexity theory, where the goal is to design protocols (i.e., rules of the debate game) that provide rigorous guarantees on correctness for judging solutions to complex problems with limited supervision. Specifically, the current best protocol has been shown to work for all problems that have sufficiently stable decompositions into subproblems. In this paper, we design a new protocol for this same class of problems that improves on the prior work in several ways. First, correctness holds in a worst-case rather than an average-case sense. Second, being honest and correct is a dominant-strategy equilibrium for both debaters, rather than a Stackelberg equilibrium. Finally, we prove black-box lower bounds, showing that our new protocol is instance-wise optimal. That is, no protocol for this class of problems can outperform ours while making only black-box queries to human judgments. We obtain these results by relating the notion of stable problem decompositions to the concept of fractional block sensitivity from query complexity.
发表机构
- UT Austin(德克萨斯大学奥斯汀分校)
- UC Berkeley(加州大学伯克利分校)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。