发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对黑箱反馈学习,发现精确动作价值函数并非方差最优,提出一种无偏且零额外成本的投影校正方法,可任意降低方差并改善学习效果。
AI 中文摘要
现代模型越来越多地通过黑箱预言机(如人类、优化求解器和外部工具)进行学习,这些预言机提供反馈但不暴露其内部机制。一种常见的补救措施是学习一个(动作)价值函数作为控制变量。在本文中,我们首先观察到,即使是一个精确的动作价值函数也可能与方差最优相差任意远。我们表明,这种差距的产生是因为价值函数最小化了每个动作自身梯度项中的噪声,而一个动作仍然可以通过共享参数影响梯度估计器的其余部分。一个简单的无偏校正,在不增加预言机成本的情况下,仍然可以将其方差减少任意大的倍数。受此启发,我们进一步证明,剩余方差可以精确地按动作分解,且无交叉项。这种分解为神经网络参数提供了一个闭式形式的方差最小化校正,该校正可以通过简单的投影计算。在实验上,我们的校正一致地减少了价值函数留下的方差,并在所有任务上改善了学习效果。所有实验的源代码可在该 https URL 获取。
英文摘要
Modern models increasingly learn through black-box oracles such as humans, optimization solvers, and external tools that provide feedback without exposing their internal mechanisms. A common remedy is to learn an (action-)value function as a control variate. In this paper, we first observe that even an exact action-value function can be arbitrarily far from variance-optimal. We show that this gap arises because the value function minimizes the noise in each action's own gradient term, while an action can still affect the rest of the gradient estimator through shared parameters. A simple unbiased correction, at no extra oracle cost, can still reduce its variance by an arbitrarily large factor. Motivated by this, we then prove that the residual variance can be decomposed exactly by actions with no cross terms. This decomposition yields a closed-form variance-minimizing correction for neural-network parameters, which can be computed by a simple projection. Empirically, our correction consistently reduces the variance left by the value function and improves learning across all tasks. The source code for all experiments is available at https://github.com/Zihao-Kevin/black_box_opt.