发表机构
CVML Lab(CVML实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CoRE 是一种用于测试时强化学习的新机制,通过基于均衡的图方法生成共识奖励,在多个基准上相比多数投票 TTRL 取得了更优性能。
AI 中文摘要
在无标签的测试数据上,强化学习缺乏真实奖励;测试时强化学习(test-time RL,TTRL)方法从模型自身的 roll-out 中推导奖励,对与 N 个采样答案的多数投票匹配的结果给予奖励。该投票机制会在正确答案为少数时将其丢弃,且对所有匹配多数的 roll-out 给予相同分数。我们用 CoRE(Consensus Rewards via Equilibrium,基于均衡的共识奖励机制)替代该投票机制:N 个 roll-out 构成一张图,其边结合了答案一致性、推理相似性和生成置信度,复制者动态(replicator dynamics)提取其主导集,从而得到精细的伪标签、每个 roll-out 的分级奖励以及每个问题的内聚性门控。CoRE 严格泛化了投票机制:多数投票是其特例;块值分析为共识在更大错误多数中恢复正确少数的情况提供了精确阈值;置信度校准可将该阈值成比例降低。在七个主干模型和五个基准(共 42 个模型-基准单元,每个单元三个随机种子)上,CoRE 使未训练的基础模型平均提升 21.7 分,而多数投票 TTRL 平均提升 20.4 分;在所有存在争议的一致场景中,CoRE 比投票机制的优势最高达 7.5 分;且 CoRE 达到投票基准平台准确率所需的步数减少了 54%至 70%。将 roll-out 组视为图而非投票箱,可在不增加额外 roll-out 成本的情况下,将脆弱的投票转化为校准后的、分级的自监督奖励。
英文摘要
On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding those that match the majority vote over $N$ sampled answers. That vote discards a correct answer whenever it is a minority and scores every majority-matching roll-out identically. We replace it with \emph{CoRE} (Consensus Rewards via Equilibrium): the $N$ roll-outs form a graph whose edges combine answer agreement, reasoning similarity, and generation confidence, and replicator dynamics extract its dominant set, yielding a refined pseudo-label, a graded per-roll-out reward, and a per-question cohesiveness gate. CoRE strictly generalizes voting: majority voting is recovered as a special case; a block-value analysis gives a sharp threshold for when consensus recovers a correct minority against a larger wrong plurality; and confidence calibration provably lowers that threshold multiplicatively. Across seven backbones and five benchmarks (42 model--benchmark cells, three seeds each), \emph{CoRE} improves the untrained base by $+21.7$ points on average versus $+20.4$ for majority-vote TTRL, wins wherever agreement is contestable with margins over the vote of up to $+7.5$ points, and reaches the voting baseline's plateau accuracy in $54$--$70$\% fewer steps. Consensus, not counting: treating the roll-out group as a graph rather than a ballot box turns a brittle vote into a calibrated, graded, self-supervised reward at no extra roll-out cost.