ICMAPE:上下文多智能体纯探索
ICMAPE: In-Context Multiagent Pure Exploration
浏览论文内容
中文总结 AI 辅助
本文提出ICMAPE,一种基于贝叶斯学习的分散式多智能体纯探索框架,通过将固定置信度识别目标转化为推理置信度奖励,联合学习集中推理网络与分散策略,在合成基准和真实数据任务中以更少探索步骤达到目标精度。
中文摘要 AI 辅助
在一些多智能体系统中,待优化的量并非外部指定的奖励,而是如主动序贯假设检验(ASHT)问题中所做的那样,获取关于环境未知属性的信息。然而,ASHT文献往往侧重于具有明确模型的有限单智能体问题,而目前缺乏能够执行主动序贯测试的实用多智能体方法。我们通过ICMAPE填补了这一空白,这是一种基于贝叶斯学习的框架,用于由推理目标驱动的分散式多智能体纯探索。ICMAPE将固定置信度识别目标转化为由推理置信度导出的奖励,从而使得标准强化学习机制能够应用于分散式纯探索。它联合学习一个集中的神经推理网络,该网络根据全局轨迹数据估计假设的后验分布,以及分散的策略,这些策略根据局部观测历史选择动作,并在达到目标置信度后学习何时停止收集数据。在两个合成基准测试和一个基于真实数据的马里兰州硝酸盐浓度监测任务上,ICMAPE-TD3以更少的探索步骤达到了目标精度。
英文摘要
In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literature tends to focus on finite single-agent problems with well-specified models, while there is currently a gap for practical multi-agent methods that can perform active sequential testing. We fill this gap with ICMAPE, a Bayesian learning-based framework for decentralized multi-agent pure-exploration driven by inference objectives. ICMAPE converts the fixed-confidence identification objective into a reward derived from inference confidence, so that standard reinforcement learning machinery can be applied to decentralized pure exploration. It jointly learns a centralized neural inference network that estimates a posterior distribution over hypotheses from global trajectory data, and decentralized policies that select actions from local observation histories and learn when to stop collecting data once the target confidence is reached. On two synthetic benchmarks and a Maryland nitrate concentration monitoring task based on real-world data, ICMAPE-TD3 achieves target accuracy with fewer exploration steps.
发表机构
- Boston University(波士顿大学)
- Broad Institute of MIT and Harvard(麻省理工学院和哈佛大学布罗德研究所)
机构由 AI 辅助整理,请以论文原文为准。