AI 中文总结
介绍用于可泛化现实世界强化学习的Building2Building基准测试,它基于EnergyPlus构建,有多样建筑配置,定义针对RL关键挑战的基准任务,能系统研究泛化与转移,推进RL研究并对HVAC控制有社会意义。
AI 中文摘要
强化学习(RL)在控制方面取得了显著成果,但学习到的策略在面对动态、动作空间、观察空间或目标的变化时仍然很脆弱,这是实际应用中的关键限制。现有基准测试的多样性和复杂性有限,难以严格研究RL中的迁移、多任务学习和元学习。我们引入了Building2Building(B2B),这是一个基于EnergyPlus构建的大规模现实供暖、通风和空调(HVAC)控制环境套件。B2B与Gymnasium接口完全兼容,具有参数化建筑生成器,可系统地生成具有异构观察和动作空间的各种建筑配置。基于此套件,我们定义了针对RL中关键开放挑战的基准任务,包括目标适应、动态适应、动作空间转移和跨域转移。通过提供具有标准化评估协议的大规模、多样化和基于物理的测试平台,B2B能够系统地研究连续控制中的泛化和转移。除了推进RL中的泛化研究之外,这个新基准测试还通过实现大规模改进的HVAC控制带来了重大的社会影响,HVAC是建筑中最耗能的系统之一。
英文摘要
Reinforcement learning (RL) has achieved strong results in control, yet learned policies remain brittle to changes in dynamics, action spaces, observation spaces, or goals, a critical limitation for real-world deployment. Existing benchmarks offer limited diversity and complexity, making it difficult to rigorously study transfer, multi-task learning, and meta-learning in RL. We introduce Building2Building (B2B), a large-scale suite of realistic Heating, Ventilation, and Air Conditioning (HVAC) control environments built on EnergyPlus, a state-of-the-art building simulator. B2B is fully compatible with the Gymnasium interface and features a parametric building generator, enabling the systematic generation of diverse building configurations with heterogeneous observation and action spaces. Based on this suite, we define benchmark tasks targeting key open challenges in RL, including goal adaptation, dynamics adaptation, action-space shifts, and cross-domain transfer. By providing a large-scale, diverse, and physically grounded testbed with standardized evaluation protocols, B2B enables systematic investigation of generalization and transfer in continuous control. Beyond advancing research on generalization in RL, this new benchmark also carries significant societal implications by enabling improved HVAC control at scale, one of the most energy-intensive systems in buildings.
Comments22 pages, 7 figures, to be published at RLC 2026