arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26391cs.LGq-bio.BM

Q-Steer:分子策略优化的动作值引导

Q-Steer: Action-Value Guidance for Molecular Policy Optimization

Xinyu Wang, Jinbo Bi, Minghu Song

首次发表
浏览论文内容

中文总结 AI 辅助

Q-Steer 是一种用于分子语言模型的生成时动作值引导原语,通过 PAVS-Q 评分器在固定在线 oracle 预算下提升分子优化性能,在所有测试组合中均取得正增益。

中文摘要 AI 辅助

受限于仅在完整分子生成后才提供奖励的分子优化场景,每一轮生成过程需做出多个局部下一个 token 的决策,这种延迟反馈机制会导致分子策略优化出现短视问题:优化器仅能得知最终分子的好坏,却无法定位到是哪些中间动作导致了该结果。我们提出 Q-Steer,这是一种用于分子语言模型的生成时动作值引导原语。Q-Steer 使用离线训练并冻结的前缀-动作值评分器 PAVS-Q,该评分器可在给定部分 SMILES 前缀时,预估选取候选下一个 token 后的下游奖励,随后将归一化的值增益添加到采样 logits 中。优化器更新规则和在线 oracle 预算保持不变,该方法的核心是在固定在线 oracle 预算下提升性能,而非要求总计算量相等。在 PMO23 数据集上,使用固定的 10000 次调用的在线预算,对两种分子语言模型主干和四种优化器开展的完整析因研究显示:Q-Steer 在所有 8 个主干-优化器组合中均提升了平均有效-唯一分数,宏观平均分数增益为 +0.033 至 +0.049,每个组合中各取得 18-20 项任务胜利。机制控制实验表明动作本身的身份至关重要:前缀广播的值几乎无影响,而打乱后的动作值会损害性能。这些结果证明 Q-Steer 是一种可复用的生成时动作值包装器,无需改变在线 oracle 预算即可提升各类优化器家族和策略主干下的平均分子优化奖励。

英文摘要

Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. This delayed-feedback interface makes molecular policy optimization myopic: an optimizer can learn that a molecule was good without knowing which intermediate actions made it good. We introduce Q-Steer, a rollout-time action-value steering primitive for molecular language models. Q-Steer uses an offline-trained and frozen prefix-action value scorer, PAVS-Q, that estimates the downstream reward of taking a candidate next token under a partial SMILES prefix, then adds a normalized value bonus to sampling logits. The optimizer update rule and online oracle budget are unchanged; the claim is fixed-online-oracle performance, not equal total compute. On PMO23 with a fixed 10,000-call online budget, complete factorial studies across two molecular language-model backbones and four optimizers show that Q-Steer improves mean valid-unique score in all eight backbone-optimizer cells, with positive macro mean-score gains between +0.033 and +0.049 and 18-20 task wins per cell. Mechanism controls show that action identity matters: prefix-broadcast values are nearly neutral, while shuffled action values harm performance. These results support Q-Steer as a reusable rollout-time action-value wrapper that improves average molecular optimization reward across optimizer families and policy backbones without changing the online oracle budget.

发表机构

  • University of Connecticut(康涅狄格大学)
  • Institute of Health and Medicine, Hefei Comprehensive National Science Center(合肥综合性国家科学中心健康与医学技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑