arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35615cs.LGcs.AI

行为基础模型用于质量多样性

Behavioral Foundation Models for Quality Diversity

Nazim Bendib, Nicolas Perrin-Gilbert, Olivier Sigaud

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出BFM-QD框架,利用行为基础模型的潜在空间进行质量多样性搜索,通过闭式无梯度策略改进算子,在连续控制基准上优于参数空间方法,尤其显著提升稀疏和欺骗性任务性能。

中文摘要 AI 辅助

行为基础模型(BFMs)是强化学习中的一种新兴范式,其作用类似于自然语言处理中的大型语言模型:它们展现出非凡的多功能性,通过利用潜在空间的结构,实现了零样本性能、快速模仿和在线适应。在这项工作中,我们研究了由BFM诱导的潜在行为空间是否可以作为有效的搜索空间,通过质量多样性(QD)方法发现大量行为多样且高性能的策略。虽然QD方法通常直接在高维策略参数空间中进行搜索,但在本文中,我们提出了BFM-QD,一个在BFM的紧凑潜在空间中进行QD搜索的框架。我们进一步表明,BFM-QD框架提供了一种闭式、无梯度的策略改进算子,该算子近似策略梯度更新,但不需要评论家训练和反向传播。在涵盖密集运动、稀疏导航和接触丰富的操作等连续控制基准测试中,BFM-QD始终优于参数空间基线,在稀疏和欺骗性设置中尤其显著,所有测试的参数空间QD方法在这些设置中性能都崩溃到接近零。这些结果表明了BFM-QD框架的有效性,得益于搜索空间的降维和来自多样化行为数据的离线预训练之间的协同作用。这将BFM定位为QD优化的通用骨干,将其效用从零样本任务解决扩展到多样化行为库的发现。

英文摘要

Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting the structure of a latent space. In this work, we investigate whether the latent behavioral space induced by BFMs can serve as an effective search space to discover large repertoires of behaviorally diverse and high-performing policies through Quality-Diversity (QD) methods. While QD methods generally search directly in high-dimensional policy parameter space, in this paper, we present BFM-QD, a framework that performs QD search in the compact latent space of a BFM. We further show that the BFM-QD framework provides a closed-form, gradient-free policy improvement operator that approximates a policy gradient update, but requires no critic training and no backpropagation. Across continuous-control benchmarks spanning dense locomotion, sparse navigation, and contact-rich manipulation, BFM-QD consistently outperforms parameter-space baselines, with particularly stark gains in sparse and deceptive settings, where all tested parameter-space QD methods collapse to near-zero performance. These results show the effectiveness of the BFM-QD framework, benefiting from the synergy between dimensionality reduction of the search space and offline pretraining from diverse behavioral data. This positions BFMs as a general-purpose backbone for QD optimization, extending their utility beyond zero-shot task solving to the discovery of diverse behavioral repertoires.

发表机构

  • Sorbonne Université(索邦大学)
  • ISIR(智能系统与机器人研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑