发表机构
University of Alicante(阿利坎特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种结合多智能体强化学习与神经进化的混合方法,利用单目视觉和紧凑网络合成群体导航控制器,在仿真中实现高效集体探索,能耗降低31.40%。
AI 中文摘要
群体机器人技术为复杂动态环境中的高级自动化提供了一种稳健且经济高效的范式,例如在搜救或环境监测中遇到的环境。该领域的一个基本挑战是数据驱动的分散控制器设计,这些控制器能够产生涌现的集体行为。本文提出了一种新颖的、由人工智能驱动的混合方法,用于自动合成群体机器人控制器,以实现自主视觉导航。该方法将多智能体强化学习与神经进化策略协同结合,特别利用交叉熵方法和协方差矩阵自适应进化策略的实现来优化预训练的单体导航策略。底层深度架构专为低成本、资源受限的平台而设计,采用仅依赖单目相机图像的紧凑神经网络。这种基于视觉的设计强调计算和能源效率,这是实际群体部署的关键要求。在高保真物理模拟器中进行的实验表明,所生成的控制器能够在多样化的室内环境中实现稳健且可扩展的集体探索。使用我们的交叉熵方法训练的控制器实现了卓越的探索覆盖率,访问的区域比协方差矩阵自适应进化策略多36.20%。至关重要的是,我们最佳的基于视觉的策略在探索性能上在统计上与依赖更昂贵距离传感器的传统方法相当,同时平均能耗显著降低了31.40%。这些发现验证了一个有效且经济可行的自主控制系统,为在现实工程应用中部署高效集体智能铺平了道路。
英文摘要
Swarm robotics presents a robust and cost-effective paradigm for advanced automation in complex, dynamic environments, such as those encountered in search and rescue or environmental monitoring. A fundamental challenge for this field is the data-driven design of decentralized controllers capable of generating emergent collective behaviors. This paper proposes a novel, AI-driven hybrid methodology for the automatic synthesis of swarm robotic controllers for autonomous visual navigation. This approach synergistically combines multi-agent reinforcement learning with neuro-evolutionary strategies, specifically leveraging implementations of the cross-entropy method and the covariance matrix adaptation evolution strategy to optimize a pre-trained individual navigation policy. The underlying deep architecture is engineered for low-cost, resource-constrained platforms, utilizing a compact neural network that relies exclusively on monocular camera imagery. This vision-based design emphasizes computational and energy efficiency, a critical requirement for practical swarm deployments. Experiments, performed in a high-fidelity physics simulator, demonstrate that the resulting controllers enable robust and scalable collective exploration of diverse indoor environments. The controller trained using our cross-entropy method achieves superior exploration coverage, visiting 36.20% more regions compared to the covariance matrix adaptation evolution strategy. Critically, our best vision-based policy achieves exploration performance statistically comparable to traditional methods relying on more expensive distance sensors, while delivering a significant 31.40% average reduction in energy consumption. These findings validate an effective and economically viable autonomous control system, establishing a path for deploying highly efficient collective intelligence in real-world engineering applications.
Comments46 pages, 21 figures. Published in Engineering Applications of Artificial Intelligence under a CC BY 4.0 license
Journal refEngineering Applications of Artificial Intelligence, Vol. 184, Part 1, 2026, 116325
DOI:10.1016/j.engappai.2026.116325