AI 中文总结
本文针对离线多任务强化学习的数据层面瓶颈,提出DaCe-DT框架,通过三种机制优化,在Meta-World上较最优方法平均提升11.73%(最优数据集)、13.34%(次优数据集),实现更优多任务性能。
AI 中文摘要
离线多任务强化学习(Offline MTRL)高度依赖预先收集数据的质量与分布。然而现有方法主要聚焦于算法优化,较少关注数据层面的改进以提升学习能力与泛化性能。本文从数据视角揭示了限制Offline MTRL性能的三个关键瓶颈:(i)在多样任务复杂度下提示长度的无效利用;(ii)随机采样的提示片段语义不相关;(iii)碎片化且不连续的轨迹所导致的误导性监督。为应对这些挑战,本文提出DaCe-DT,这是一个对异构任务复杂度与数据质量不敏感的鲁棒离线MTRL框架,包含长度门控提示掩码(LGPM)、检索增强的提示构建(RAPC)以及值自适应回报校准(VARC)。这些机制共同使DaCe-DT实现以数据为中心的提示适配与轨迹优化,在异构离线数据与任务中实现鲁棒的多任务泛化与稳定的策略学习。在Meta-World上的实验结果显示,DaCe-DT始终优于现有最优方法,在最优数据集上平均提升11.73%,在次优数据集上提升13.34%,证明其从不完美数据中稳定学习、提升整体多任务性能的有效性。
英文摘要
Offline multi-task reinforcement learning (Offline MTRL) heavily depends on the quality and distribution of pre-collected data. However, existing methods mainly focus on algorithmic optimization, with less emphasis on data-level improvements to enhance learning ability and generalization performance. This paper, from a data perspective, reveals three key bottlenecks that limit Offline MTRL performance:(i) ineffective utilization of prompts length under diverse task complexities, and (ii) semantic irrelevance of randomly sampled prompt segments, (iii) misleading supervision induced by fragmented and discontinuous trajectories. To address these challenges, we propose DaCe-DT, a robust offline MTRL framework designed to be insensitive to heterogeneous task complexities and data quality, featuring length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC). Together, these mechanisms enable DaCe-DT to deliver data-centric prompt adaptation and trajectory refinement, resulting in robust multi-task generalization and stable policy learning amid heterogeneous offline data and tasks. Experimental results on Meta-World show that DaCe-DT consistently outperforms state-of-the-art methods, achieving an average improvement of 11.73% on optimal datasets and an improvement of 13.34% on suboptimal datasets, demonstrating its effectiveness in learning stably from imperfect data and improving overall multi-task performance.
Comments25 pages,NeurIPS-2026