arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38065cs.LGcs.AI

Jaxolotl:基于LTL的多任务强化学习统一高性能基准套件

Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Multi-Task RL

Mathias Jackermeier, Jacques Cloete, Alessandro Abate

首次发表
浏览论文内容

中文总结 AI 辅助

针对多任务LTL-RL方法难以比较且计算成本高的问题,提出统一高性能基准套件Jaxolotl,通过JAX实现和预编译实现220倍加速,系统评估揭示方法间互补优劣。

中文摘要 AI 辅助

训练智能体遵循任意指令是多任务强化学习(RL)的一个重要目标。线性时序逻辑(LTL)为向智能体指定指令提供了一种精确且结构化的形式体系,并已被成功用于训练通用多任务策略。然而,现有方法在实现、任务分布和评估协议上的差异使得它们难以比较,而高昂的计算成本限制了实验的规模和统计可靠性。我们引入了Jaxolotl,一个用于多任务LTL-RL的统一高性能基准套件,以解决这些问题。Jaxolotl提供了六个代表性算法和四个环境的模块化、端到端JAX实现,以及新整理的任务套件和标准化、统计稳健的评估协议。通过将符号任务表示预编译为静态数组,Jaxolotl实现了完全JIT编译的训练和评估,实现了高达220倍的端到端加速,并支持在显著更大的实验规模下进行受控比较。我们利用该框架系统评估了现有方法,揭示了互补的优势和局限:能够进行非短视推理的通用方法随着命题数量的增长而表现挣扎,而扩展性更强的方法依赖于环境特定的假设并遭受短视问题。

英文摘要

Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted for training generalist multi-task policies. However, differences in implementations, task distributions, and evaluation protocols make existing methods difficult to compare, while high computational costs limit the scale and statistical reliability of experiments. We introduce Jaxolotl, a unified high-performance benchmark suite for multi-task LTL-RL to address these concerns. Jaxolotl provides a modular, end-to-end JAX implementation of six representative algorithms and four environments, together with newly curated task suites and a standardised, statistically robust evaluation protocol. By precompiling symbolic task representations into static arrays, Jaxolotl enables fully JIT-compiled training and evaluation, achieving end-to-end speedups of up to $220\times$ and supporting controlled comparisons at substantially greater experimental scale. We use this framework to systematically evaluate existing approaches, revealing complementary strengths and limitations: general methods capable of non-myopic reasoning struggle as the number of propositions grows, while methods with stronger scaling rely on environment-specific assumptions and suffer from myopia.

发表机构

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

↑