arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01652cs.LGcs.AI

通过语义回滚分析的迭代式策略精化

Iterative Policy Refinement through Semantic Rollout Analysis

Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons

首次发表
浏览论文内容

中文总结 AI 辅助

提出闭环框架,利用LLM分析策略回滚的表格数据,迭代精化结构化策略,无需人工指令,在赛车和开门任务上性能提升最高15%,计算量减少75%。

中文摘要 AI 辅助

结构化策略通过在模仿学习中引入任务特定的归纳偏置,提高了效率、鲁棒性和可解释性,但现有的结构生成方法要么依赖大量人工输入,要么依赖编码在大型语言模型中的静态领域知识,这些知识可能与专家演示不一致。我们提出一个闭环框架,利用大型语言模型引导的策略回滚分析来迭代地精化结构化策略。通过将回滚记录为语义上有意义的表格数据,并提示大型语言模型生成诊断分析代码,我们的方法识别策略结构中的次优之处,并在无需人工指令的情况下迭代修正它们。在赛车和开门任务上的实验表明,我们的方法相比零样本大型语言模型生成的结构,将模仿学习性能提升了最多15%,并且达到相同的强化学习性能所需计算量减少了75%。这些结果证明,表格回滚分析提供了有效的反馈信号,能将大型语言模型生成的策略结构与专家演示对齐,并且我们可以利用它自动生成良好的策略结构。

英文摘要

Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • Meta

机构由 AI 辅助整理,请以论文原文为准。

↑