arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21897cs.AI

从求解器反馈到可信规划:用于符号规划的多角色强化学习

From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning

Chenghao Zhang, Yikai Mao, Haoyu Gao, Saisai Hu, Yuxi Cheng, Shanqi Liu, Dan Roth

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出基于求解器的多角色强化学习框架,仅用求解器反馈实现无人工演示的自然语言转PDDL,在PlanBench上显著提升规划成功率与可信性,降低语义漂移。

中文摘要 AI 辅助

可靠的规划需要将自然语言指令转换为可执行的符号规范,但大型语言模型在没有昂贵的PDDL注释的情况下仍表现脆弱,且可能以语义不可靠的方式利用求解器的成功。我们研究如何仅使用求解器反馈学习可信的自然语言到PDDL的形式化,无需人工编写的演示。我们提出了基于求解器的多角色强化学习框架,其中单个语言模型充当生成、验证和修复的Actor、Judge和Editor角色:Actor提出PDDL规范,Judge提供求解器校准的质量信号,Editor执行有界的诊断条件细化。在PlanBench上,我们的方法将LLM+P的平均成功率从35.5%提升至70.8%,达到66.3%的可信成功率,并将语义漂移降低至6.4%。这些结果表明,将求解器反馈组织为生成、验证和修复角色,可实现更具可扩展性和可信的无注释符号规划。

英文摘要

Reliable planning requires converting natural-language instructions into executable symbolic specifications, yet large language models remain brittle without costly PDDL annotations and may exploit solver success in semantically unfaithful ways. We study how to learn faithful natural-language-to-PDDL formalization using only solver feedback, without human-written demonstrations. We propose a solvergrounded multi-role reinforcement learning framework where a single language model acts as an Actor, Judge, and Editor for generation, verification, and repair. The Actor proposes PDDL specifications, the Judge provides a solver-calibrated quality signal, and the Editor performs bounded diagnostic-conditioned refinement. On PlanBench, our method improves average success from 35.5% for LLM+P to 70.8%, achieves 66.3% faithful success, and reduces semantic drift to 6.4%. These results show that organizing solver feedback into generation, verification, and repair roles enables more scalable and faithful annotation-free symbolic planning

发表机构

  • University of Pennsylvania(宾夕法尼亚大学)
  • Georgia Institute of Technology(佐治亚理工学院)
  • Pace University(佩斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑