发表机构
University of Pennsylvania; Shanghai AI Laboratory(宾夕法尼亚大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出最小见证强化学习(MWRL),通过基于覆盖贡献的信用分配,从黑盒验证器反馈中恢复多个最小充分见证,扩展强化学习至多解识别。
AI 中文摘要
“产生某个结果所需的不可约条件是什么?”是计算和科学中反复出现的最常见问题之一。其答案,即最小充分见证,正是我们所说的解释、机制和原因。这些问题通常要求多个最小见证,但标准的强化学习方法可能只揭示一个解决方案或冗余的解决方案。我们将此问题形式化为最小见证识别,并引入最小见证强化学习(MWRL)。MWRL 将策略采样的成功提案所认证的集合取并集,并依据每个提案对群体并集在缺少该提案时会损失的覆盖范围来赋予其信用。这种信用分配直接源于问题定义,统一了从单一黑盒验证器比特中实现最小性和恢复替代方案的需求。在此原则下,我们推导出一个值迭代规划器,可恢复全部见证家族,以及一个可扩展至大型语言模型的策略梯度方法。在不同实验设置中,MWRL 能恢复大多数最小见证,而其他方法则返回冗余超集或单一见证。通过使见证家族可从验证器反馈中学习,MWRL 将强化学习的范围扩展到单解优化之外。我们的代码可从此 https URL 获取。
英文摘要
``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by explanations, mechanisms and reasons. These problems usually ask for multiple minimal witnesses, yet standard RL methods may reveal only one solution or redundant ones. We formalize this problem as minimal-witness identification and introduce Minimal-Witness Reinforcement Learning (MWRL). MWRL takes the union of the sets certified by successful proposals sampled from the policy and credits each proposal for the coverage the group union would lose without that proposal. This credit assignment, derived directly from the problem definition, unifies the demands for minimality and recovery of alternatives from a single black-box verifier bit. Under this principle, we derive a value iteration planner that recovers the entire family of witnesses and a policy gradient method that can scale to large language models. Across different experimental settings, MWRL recovers most minimal witnesses, while other methods return redundant supersets or a single witness. By making witness families learnable from verifier feedback, MWRL expands the scope of reinforcement learning beyond single-solution optimization. Our code is available at https://github.com/TSUITUENYUE/MWRL.