arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CDRL:用于中微子味模型发现的认证驱动强化学习

CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

Piyush Jha, Jake Rudolph, Victoria Knapp-Pérez, Max Fieg, Aishik Ghosh, Vijay Ganesh

arXiv 2608.20686首次发表:更新:

发表机构

School of Physics, Georgia Institute of Technology; University of California, Irvine; Halluminate; Fermilab; Lawrence Berkeley National Laboratory(佐治亚理工学院物理学院; 加利福尼亚大学欧文分校; 哈卢米奈特公司; 费米国家加速器实验室; 劳伦斯伯克利国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出CDRL框架,利用符号推理工具的结构化反馈将失败证书转为可重用约束,在中微子味模型发现任务上,较现有RL方法大幅提升有效模型与中微子模型发现率,减少候选评估量,还可提取可解释规则作为软约束进一步提升性能。

AI 中文摘要

许多科学发现问题需要在复杂的领域约束下搜索组合假设空间。强化学习(RL)提供了一种有前景的方法,但现有方法依赖标量奖励,其关于候选解失败原因的信息有限,导致智能体反复探索无效区域。我们引入认证驱动强化学习(CDRL),这是一个利用符号推理工具的结构化反馈的框架。当候选解违反领域约束时,这些工具会生成证书,标识导致失败的动作。CDRL将这些证书转换为可重用的约束,消除各类无效解,并引导探索走向有效区域。我们在理论粒子物理学的中微子味模型发现任务上评估CDRL,该任务的假设空间超过$10^{26}$个可能模型,并将其与之前用于该任务的最先进RL方法进行比较。在三个理论空间中,CDRL实现了高达1.95倍的有效模型率、高达6.33倍的中微子模型率,同时评估的候选数量减少了多达4倍。我们进一步使用事后决策树框架从搜索轨迹中提取40条可解释规则,并显示将这些规则作为软约束重用,在所有三个理论空间中可使有效模型率提升高达2倍,中微子模型发现率提升高达3倍。这些结果表明,CDRL能在组合搜索空间中发现可重用结构,并为科学模型发现提供了通用框架。

英文摘要

Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints. Reinforcement learning (RL) offers a promising approach, but existing methods rely on scalar rewards that provide limited information about why candidate solutions fail, leading agents to repeatedly explore invalid regions. We introduce Certification-Driven Reinforcement Learning (CDRL), a framework that leverages structured feedback from symbolic reasoning tools. When a candidate violates domain constraints, these tools produce certificates identifying the actions responsible for failure. CDRL converts these certificates into reusable constraints that eliminate classes of invalid solutions and guide exploration toward valid regions. We evaluate CDRL on neutrino flavor model discovery in theoretical particle physics, where the hypothesis space exceeds $10^{26}$ possible models, and compare it with the state-of-the-art RL approach previously used for this task. Across three theory spaces, CDRL achieves up to 1.95$\times$ higher valid model rates and up to 6.33$\times$ higher neutrino model rates while evaluating up to 4$\times$ fewer candidates. We further extract 40 interpretable rules from search trajectories using a post-hoc decision-tree framework and show that reusing them as soft constraints yields gains of up to 2$\times$ in valid model rates and 3$\times$ in neutrino model discovery across all three theory spaces. These results suggest that CDRL uncovers reusable structure in combinatorial search spaces and provides a general framework for scientific model discovery.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑