arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01929cs.CE

面向条件生成流网络(GFlowNets)的风险敏感奖励组合

Risk-Sensitive Reward Composition for Conditional GFlowNets

Carine Ribeiro dos Santos, Ina Pöhner

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对柔性蛋白构象的药物设计问题,提出结合CVaR与歧义半径的风险敏感奖励组合方法,训练单个条件GFlowNet,可高效找到通过所有构象的优质候选分子,验证了其性能优势。

中文摘要 AI 辅助

基于结构的药物设计中,生成流网络(GFlowNets)以一个刚性蛋白质结构为条件。柔性靶标具有多种不同的结构形状(即构象),每种构象在模拟时间中占据一定比例。针对所有构象对候选分子进行评分时,会出现一个开放性问题:如何将K个评分合并为一个奖励?设计者不能任意选择,因为种群存在模拟误差,且生物学特性决定了哪些构象是关键约束——未通过某一构象的候选分子会被直接淘汰,而非仅排名降低,现有标准规则无法处理该问题。我们将奖励由条件风险价值(CVaR,一种最坏情况评分规则)和表示对给定权重不信任度的歧义半径组合而成,共同定义了一类目标,可由单个条件GFlowNet进行摊销。我们探究此类采样器是否可在完全可枚举的合成世界中训练(其中所有误差为精确值而非估计值)。与平均策略相比,尾部定价策略可将2至10倍的质量转移至通过所有构象的候选分子。单个网络覆盖该类目标的误差在完美采样器误差的0.37至2.7倍范围内。精确KL(KL散度)神谕(即基于真实目标训练的副本)可判断性能不足源于优化器还是架构。当优质候选分子稀少时,探索起关键作用:注入未见过的状态可找到0.987至1.000的优质区域,而对已访问状态重新加权仅能找到0.35至0.76的优质区域。

英文摘要

Generative Flow Networks (GFlowNets) for structure-based drug design condition on one rigid protein structure. A flexible target holds several distinct structural shapes, its conformations, each occupied for a fraction of the simulation time. Scoring a candidate against all of them raises an open question: how do K scores become one reward? The designer cannot choose arbitrarily. Populations carry simulation error, and biology dictates which conformations are deal-breakers, so a candidate that fails one is disqualified, not merely ranked lower. No standard rule captures this. We compose the reward from a conditional value-at-risk (CVaR), a worst-case score rule, and an ambiguity radius expressing distrust in the stated weights. Together, these define a family of targets, amortised by a single conditional GFlowNet. We answer whether such a sampler can be trained on fully enumerable synthetic worlds, where every error is exact rather than estimated. Pricing the tail rather than averaging moves 2-10 times more mass to candidates that pass every conformation. One network covers the family to within 0.37-2.7x the error of a perfect sampler. An exact-KL oracle, a copy trained on the true target, shows if a shortfall is the optimiser's or the architecture's. When good candidates are rare, exploration decides: injecting unseen states finds 0.987-1.000 of good regions, while reweighting visited finds 0.35-0.76.

发表机构

  • Instituto de Química, Departamento de Química Orgânica, Universidade Federal do Rio de Janeiro(里约热内卢联邦大学化学系有机化学研究所)
  • School of Pharmacy, University of Eastern Finland(东芬兰大学药学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑