基于正确示范的多尺度奖励对冲
Multiscale Reward Hedging from Correct Demonstrations
浏览论文内容
中文总结 AI 辅助
该研究针对从正确示范学习的问题,提出多尺度奖励对冲方法,获首个无时间范围的多项式有限界,在上下文推荐等任务中取得良好效果,且计算效率在特定场景下可优化。
中文摘要 AI 辅助
当存在多个正确答案时,从正确示范中学习比监督学习更困难:学习者做出预测后,仅能看到一个有效答案,无法判断自身答案是否有效,也无法获得任何奖励。现有奖励对冲的保证因此假设奖励类是有限的。我们针对连续类给出了首个无时间范围(horizon-free)的保证,核心是在每个精度尺度上对容忍最优性测试进行一次共享投票对冲。目标奖励在每个尺度都有一个存活代理,且与该尺度的差距超过一定值的预测会使代理翻倍。这产生了同时尾界:$|\{t:\ell_t>2^{-j}\}|\leq \log_2\mathcal N(\mathcal G,2^{-j-1})+j$,其中$\mathcal G$是最优性差距函数类。对尾界积分可得到累积隐藏差距,其受度量熵积分约束,与回合数无关。多项式熵$(A/\epsilon)^d$给出总差距为$O(d\log A)$和快速统计速率$O(d/m)$。对于有界线性上下文推荐,该结果对任意紧凑菜单给出$O(d)$的遗憾。这是首个对菜单无结构限制的多项式有限界,代价是采用非恰当预测。尽管通用投票计算成本较高,但对于一维利普希茨参数曲线,它恰好是多项式时间的。固定半径二阶推荐对大小为$K$的菜单耗时$O(KT^2)$。我们还证明了$\Omega(d)$下界、低秩和有界ReLU网络推论,以及仅增加示范者累积次优性的鲁棒定理。可复现的自适应压力测试说明了预测尺度适应。分解后,精确的MovieLens审计在10个用户上耗时1.7 CPU秒,且比示范评分策略和恰当在线基线的平均潜在差距更小。学习者仅使用动作示范,从不观察奖励或损失。
英文摘要
Learning from correct demonstrations is harder than supervised learning when many answers are correct: after predicting, the learner sees one valid answer but not whether its own answer was valid, nor any reward. Existing reward-hedging guarantees consequently assume a finite reward class. We give the first horizon-free guarantee for continuous classes. The key is to hedge in one shared vote over tolerant optimality tests at every accuracy scale. A target reward has one surviving proxy per scale, and a prediction with gap above that scale doubles the proxy. This yields the simultaneous tail bound $|\{t:\ell_t>2^{-j}\}|\leq \log_2\mathcal N(\mathcal G,2^{-j-1})+j$, where $\mathcal G$ is the class of optimality-gap functions. Integrating the tails gives cumulative hidden gap bounded by a metric-entropy integral, independently of the number of rounds. Polynomial entropy $(A/ε)^d$ gives $O(d\log A)$ total gap and a fast $O(d/m)$ statistical rate. For bounded linear contextual recommendation, the result is $O(d)$ regret for arbitrary compact menus. This is the first polynomial finite bound without structural restrictions on the menus, at the price of improper prediction. Although the general vote can be expensive, it is exactly polynomial-time for one-dimensional Lipschitz parameter curves. Fixed-radius rank-two recommendation takes $O(KT^2)$ time for menus of size $K$. We also prove an $Ω(d)$ lower bound, low-rank and bounded ReLU-network corollaries, and a robust theorem that adds only the demonstrator's cumulative suboptimality. A reproducible adaptive stress test illustrates the predicted scale adaptation. After factorization, an exact MovieLens audit runs in 1.7 CPU seconds across ten users and improves mean latent gap over both a demonstrated-rating policy and a proper online baseline. The learner uses only action demonstrations and never observes a reward or a loss.
发表机构
- Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。