arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

好的排序器,差的目标:表达性策略搜索下的双线性对比评判器

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

Ayushman Singh, Siddharth Aphale

arXiv 2607.27422首次发表:更新:

发表机构

Stanford University; Sesame AI(斯坦福大学; 芝麻AI公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现双线性对比评判器的排序能力与其优化安全性不匹配,其存在范数漂移等问题,在多任务中无法可靠排序动作,需结合价值校准标量用于动作选择。

AI 中文摘要

良好的动作排序并不能保证对比评判器可被安全最大化,这类评判器正越来越多地用作K选最优、规划及评判器引导生成的类值目标。无界双线性得分会让大嵌入范数放大非支持值,而余弦约束无法消除失效情况;受控支持分解将大部分原始双线性遗憾归因于范数漂移。不过余弦及混合评判器仍会从多数动作池中选择非支持动作,产生的遗憾也相当。在四个OGBench导航任务中,前10%得分的对比得分校准性弱或反转,且无法按价值对固定查询动作排序;经贝尔曼训练的TD-Q则表现出色,包括在参数匹配的函数类控制场景中。实际成本取决于任务:模拟器回滚显示PointMaze和精确Q*玩具上的单步选择成本,而AntMaze和HumanoidMaze上存在效力充足的零值,控制器可自我修正。训练/读出分解将排序丢失归因于余弦训练目标,原始训练的嵌入在推理时归一化后仍保留弱排序。因此候选最大化可利用范数漂移、得分饱和或支持内排序错误导致的假阳性;对比评判器在导航和操作任务上仍可用作兼容性排序器,但动作选择需经价值校准的标量。

英文摘要

Good action rankings do not make a contrastive critic safe to maximize. These critics increasingly act as value-like objectives for best-of-$K$ selection, planning, and critic-guided generation. Unbounded bilinear scores can let large embedding norms inflate off-support values, but cosine bounding does not remove the failure. A controlled support decomposition attributes most raw bilinear regret to norm drift. Cosine and hybrid critics nevertheless select off-support actions from most pools and incur comparable regret. Contrastive scores are weakly calibrated or inverted in the top score decile across four OGBench navigation tasks, and they fail to order fixed-query actions by value. Bellman-trained TD-Q succeeds, including in a parameter-matched function-class control. Realized costs depend on the task: simulator rollouts reveal single-step selection costs on PointMaze and the exact-$Q^*$ toy but well-powered nulls on AntMaze and HumanoidMaze, where the controller can self-correct. A training/readout decomposition traces the lost ordering to the cosine training objective; raw-trained embeddings retain weak ordering after inference-time normalization. Candidate maximization can therefore exploit false positives caused by norm drift, score saturation, or in-support misranking. Contrastive critics remain useful compatibility rankers on navigation and manipulation tasks, but action selection requires a value-calibrated scalar.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑