arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ROOT:为用户指定的具身行为发现奖励函数

ROOT: Discovering Rewards for User-Specified Embodied Behaviors

Eren Sadikoglu, Aditya Taparia, Xinyuan Liu, Ransalu Senanayake

arXiv 2610.04250首次发表:更新:

发表机构

Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ROOT通过可观察树搜索奖励函数,利用视频语言模型诊断行为失败,在四个具身实体上实现高达86.8%的运动完整性和16.5%的行为对齐提升,优于现有LLM方法。

AI 中文摘要

具身控制的强化学习仍然受到奖励函数指定难度的制约。尽管最近基于大型语言模型(LLM)的方法可以从自然语言描述中合成奖励函数,但它们往往无法捕捉人类所关心的微妙行为属性,例如自然步态、姿态和运动风格。这一局限性源于许多期望行为更容易通过视觉识别而非编码为奖励函数。我们提出了奖励优化通过可观察树(ROOT),一个用于发现奖励函数的框架,该框架使学习到的策略与用户指定的具身行为对齐。ROOT不单纯依赖标量训练统计,而是将奖励设计视为在持久实验树上的观察引导搜索,该树存储奖励程序、训练策略和 rollout 观察,并结合由视频语言模型提炼的行为洞察,以诊断行为失败并指导后续奖励改进。我们在四个具身实体上的七个任务中评估了 ROOT:模拟的 Hopper、HalfCheetah、Ant、Unitree Go2,以及真实世界的 Unitree Go2。ROOT产生的行为比现有基于LLM的奖励生成方法更符合用户意图,实现了高达86.8%的运动完整性准确率,并将Vid-LLM行为对齐从3.56/5提高到4.14/5,相比基线提升了16.5%。人类评估进一步支持这些结果,ROOT在51-63%的成对比较中被优先选择。

英文摘要

Reinforcement learning for embodied control remains constrained by the difficulty of reward specification. Although recent large language model (LLM)-based methods can synthesize reward functions from natural-language descriptions, they often fail to capture subtle behavioral properties that humans care about, such as natural gait, posture, and movement style. This limitation arises because many desired behaviors are easier to recognize visually than to encode in a reward function. We introduce Reward Optimization via Observable Trees (ROOT), a framework for discovering reward functions that align learned policies with user-specified embodied behaviors. Rather than relying solely on scalar training statistics, ROOT casts reward design as an observation-guided search over a persistent experiment tree that stores reward programs, trained policies, and rollout observations, together with behavioral insights distilled by a video-language model, to diagnose behavioral failures and guide subsequent reward refinements. We evaluate ROOT on seven tasks across four embodiments: simulated Hopper, HalfCheetah, Ant, Unitree Go2, and as well as the real-world Unitree Go2. ROOT produces behaviors that better align with user intent than those generated by existing LLM-based reward-generation methods, achieving up to 86.8% locomotion-completeness accuracy and improving Vid-LLM behavioral alignment from 3.56/5 to 4.14/5, a 16.5% improvement over baselines. Human evaluations further support these results, with ROOT preferred in 51-63% of pairwise comparisons.

Comments13 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑