多次响应的统计优势:从示范中学习的视角
The Statistical Benefits of Multiple Responses for Learning from Demonstrations
- University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究非最优示范下多次响应的统计优势,发现 pass@$k$ 可将最坏情况精度依赖从 $1/\varepsilon^2$ 降至 $1/\varepsilon$,并在标准评估下改进奖励类别依赖,提供匹配上下界及贪心算法。
AI中文摘要:
许多生成系统会返回多个候选响应,并根据最佳响应进行评估。近期研究表明,当示范为最优时,pass@$k$ 可以将从示范中学习的样本复杂度降低一个 $k$ 的对数因子。我们探讨当示范者并非最优时会发生什么。我们发现,在此设定下,多次响应提供了质量上更强的优势。在一个无奖励反馈的有限奖励类别模型中,从 pass@$1$ 转向任意 $k\ge2$ 的 pass@$k$,会将目标精度的最坏情况依赖从 $1/\varepsilon^2$ 改变为 $1/\varepsilon$,且对示范者质量均匀成立。在标准评估下,即训练前固定一个未知奖励,增加 $k$ 提供了额外且不同的优势:对大小为 $N$ 的奖励类别的最优依赖从 $\log N$ 改进为 $\log N/\log k$。我们进一步表明这两种效应可以分离。在稳健评估下,即一个学习到的策略必须同时与示范者竞争类别中的每个奖励,快速的 $1/\varepsilon$ 依赖持续存在,而 $1/\log k$ 的改进可能消失。我们在相应机制中建立了匹配的上界和下界,并给出了一个贪心乘法权重学习器,该学习器在不对示范者质量做任何假设的情况下达到上界。
英文摘要:
Many generative systems return multiple candidate responses and are evaluated according to the best one. Recent work shows that, when demonstrations are optimal, pass@$k$ can reduce the sample complexity of learning from demonstrations by a logarithmic factor in $k$. We ask what happens when the demonstrator is not assumed to be optimal. We find that multiple responses provide a qualitatively stronger benefit in this setting. In a finite reward-class model with no reward feedback, moving from pass@$1$ to any pass@$k$ with $k\ge2$ changes the worst-case dependence on target accuracy from $1/\varepsilon^2$ to $1/\varepsilon$, uniformly over demonstrator quality. Under standard evaluation, where an unknown reward is fixed before training, increasing $k$ provides an additional and distinct benefit: the optimal dependence on a reward class of size $N$ improves from $\log N$ to $\log N/\log k$. We further show that these two effects can be separated. Under robust evaluation, where one learned policy must compete with the demonstrator simultaneously for every reward in the class, the fast $1/\varepsilon$ dependence persists, while the $1/\log k$ improvement can disappear. We establish matching upper and lower bounds in the corresponding regimes and give a greedy multiplicative-weights learner achieving the upper bounds without any assumption on demonstrator quality.