发表机构
Engineering Research Center of Machine Learning and Industry Intelligence; National University of Singapore; University of Leeds; Sichuan University(机器学习与产业智能工程研究中心; 新加坡国立大学; 利兹大学; 四川大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM智能体主动探索能力的两大瓶颈,提出含探索性数据构建、对比信号引导的RL优化的方法,经实验验证其有效性并公开代码。
AI 中文摘要
我们研究大语言模型智能体(LLM Agents)的主动探索能力,即智能体探索环境以获取信息、优化未来决策的能力。在此方面,我们首先确定了阻碍该能力的两个核心瓶颈,随后提出\textbf{Ours}这一旨在注入并优化主动探索能力的新方法。具体而言,\textbf{Ours}包含两个组件:(1)探索性数据构建,即合成富含探索内容的轨迹,以缓解标准示范的后视偏差;(2)对比信号引导的强化学习(RL)优化,即利用对比轨迹对区分有益探索与冗余漫游。大量实验证明了\textbf{Ours}的有效性,并为主动探索的特征提供了见解,代码可访问该链接:this https URL。
英文摘要
We study proactive exploration in LLM agents, i.e., the ability to explore an environment to acquire information that improves future decision-making. In this regard, we first identify two fundamental bottlenecks that hinder this capability and then propose \ours, a novel method designed to instill and refine proactive exploration. Specifically, \ours\ consists of two components: (1) Exploratory Data Construction, which synthesizes exploration-rich trajectories to mitigate the hindsight bias of standard demonstrations; and (2) RL Optimization with Contrastive Signal Guidance, which leverages contrastive trajectory pairs to distinguish productive exploration from redundant wandering. Extensive experiments demonstrate the effectiveness of \ours\ and provide insights into the characteristics of proactive exploration. Our code is available at: https://github.com/GuanZhizhao/SAFARI.