arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2411.13072cs.ROcs.AI

AMaze:用于可泛化智能体快速原型设计的直观基准生成器

AMaze: An intuitive benchmark generator for fast prototyping of generalizable agents

Kevin Godin-Dubois, Karine Miras, Anna V. Kononova

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出 AMaze 基准生成器,通过可定制、具欺骗性的视觉迷宫支持具身智能体泛化训练与人在回路分析;实验显示脚手架式和交互式训练的泛化增益达50%–100%。

中文摘要 AI 辅助

训练智能体的传统方法通常涉及一个复杂度极低的单一确定性环境,以解决机器人运动或计算机视觉等各种任务。然而,在静态环境中训练的智能体缺乏泛化能力,限制了其在更广泛场景中的潜力。因此,近期的基准经常依赖多个环境,例如通过提供随机噪声、简单排列或完全不同的设置。在实践中,此类环境集合主要来自成本高昂的人工设计过程,或对随机数生成器的不加约束的使用。在本工作中,我们介绍了 AMaze,这是一种新型基准生成器,其中具身智能体必须通过解读具有任意复杂度和欺骗性的视觉标志来在迷宫中导航。该生成器通过轻松生成特定特征的迷宫以及直观理解所得智能体的策略,促进人机交互。作为概念验证,我们在一个欺骗性有限的简单、完全离散案例中展示了该生成器的能力。智能体在三种不同机制下训练(单次、脚手架式、交互式),结果表明后两种情况在泛化能力方面优于直接训练。事实上,根据泛化指标、训练机制和算法的组合,中位增益在50%至100%之间,并且通过交互式训练实现了最高性能,从而证明了可控的人在回路基准生成器的优势。

英文摘要

Traditional approaches to training agents have generally involved a single, deterministic environment of minimal complexity to solve various tasks such as robot locomotion or computer vision. However, agents trained in static environments lack generalization capabilities, limiting their potential in broader scenarios. Thus, recent benchmarks frequently rely on multiple environments, for instance, by providing stochastic noise, simple permutations, or altogether different settings. In practice, such collections result mainly from costly human-designed processes or the liberal use of random number generators. In this work, we introduce AMaze, a novel benchmark generator in which embodied agents must navigate a maze by interpreting visual signs of arbitrary complexities and deceptiveness. This generator promotes human interaction through the easy generation of feature-specific mazes and an intuitive understanding of the resulting agents' strategies. As a proof-of-concept, we demonstrate the capabilities of the generator in a simple, fully discrete case with limited deceptiveness. Agents were trained under three different regimes (one-shot, scaffolding, interactive), and the results showed that the latter two cases outperform direct training in terms of generalization capabilities. Indeed, depending on the combination of generalization metric, training regime, and algorithm, the median gain ranged from 50% to 100% and maximal performance was achieved through interactive training, thereby demonstrating the benefits of a controllable human-in-the-loop benchmark generator.

补充信息

↑