发表机构
Starling Research Institute(斯塔林研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种强化学习智能体,将代数建模为带动态动作空间的MDP,在符号方程求解任务上优于非学习型搜索方法,尤其在指数方程上表现突出,但仅针对四类受控开式方程有效。
AI 中文摘要
我们提出了一种逐步求解符号方程的强化学习智能体,覆盖非线性闭式方程(根式、指数、三角函数)以及需要变量代换(CoV,如配方法)的受控受限开式族。我们将代数问题建模为具有动态动作空间和树状策略(TreeMLP)的马尔可夫决策过程(MDP)。主策略仅从奖励中学习,无需监督求解轨迹;变量代换(CoV)替换来自可与计算机代数系统(CAS)调用互换的监督生成器。在闭式方程上,该智能体在单一策略下达到了CommonCore数据集上的现有最佳性能(0.93 贪心搜索,优于ConPoLe的0.925)。在四个手动设计的受限开式族(二次、三次、四次、指数)上,它达到了0.79 束搜索/0.67 贪心搜索的准确率,超过了最强的非学习型搜索方法A*(0.64)。学习到的变量代换(CoV)时机仅在指数族上有实际意义,该族需要嵌套变量代换,其中自然规则无法求解任何保留的方程,而该策略仅通过奖励就求解了75%的方程。在10倍规模下,出现了明显的种子级双峰分布;UCB学习进度课程显示出缓解该双峰分布的非显著正向趋势。我们不声称能求解一般开式方程:所有开式方程结果均限于这四个受控族。
英文摘要
We present a reinforcement-learning agent that solves symbolic equations step by step, covering both nonlinear closed equations (radicals, exponentials, trigonometric) and a controlled class of restricted-open families requiring a change of variables (CoV) such as completing the square. We cast algebra as an MDP with a dynamic action space and a tree-structured policy (TreeMLP). The main policy learns from reward alone with no supervised solution traces; the CoV substitution comes from a supervised generator interchangeable with a CAS call. On closed equations the agent matches the prior best on CommonCore (0.93 greedy vs. ConPoLe's 0.925) under a single policy. On four hand-designed restricted-open families (quadratic, cubic, quartic, exponential) it reaches 0.79 beam / 0.67 greedy, exceeding the strongest non-learned search (A-star, 0.64). Learned CoV timing has content only on the exponential family, the one requiring a nested CoV, where a natural rule solves none of the held-out equations while the policy solves 75% from reward alone. At 10x scale a sharp seed-level bimodality emerges; a UCB learning-progress curriculum shows a non-significant positive trend toward mitigating it. We do not claim general open-equation solving: every open-equation result is confined to these four controlled families.