arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于强化学习的时序逻辑引导通用任务表示

Temporal Logic Guided Universal Task Representations for Reinforcement Learning

Hao Zhang, Zhangli Zhou, Zhen Kan

arXiv 2608.15509首次发表:更新:

发表机构

University of Science and Technology of China; Anhui Provincial Key Laboratory of Humanoid Robots; Chinese Academy of Sciences(中国科学技术大学; 安徽省人形机器人重点实验室; 中国科学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出受时序逻辑启发的通用任务表示框架LOTUS,将其集成至强化学习算法,通过LTL编码器建模语义、双模拟度量保障稳定性,在多场景下的学习效率、泛化能力等指标优于现有方法。

AI 中文摘要

受任务引导的智能体在各类复杂任务中展现出优异性能,但现有多数任务表示算法针对特定场景设计,难以在多样场景中泛化;且它们通常依赖强化学习控制器的梯度信号更新权重,会降低表示质量与学习效率。为克服这些局限,本文提出LOTUS——一种受时序逻辑启发的通用任务表示框架,可无缝集成至任意强化学习(RL)算法,以提升智能体在多样任务设置下的性能。具体而言,本文设计了一种新型任务表示架构,能够从线性时序逻辑(LTL)公式中建模关系并提取任务语义;还引入了更高效的更新机制,将LTL编码器视为策略,从而提升表示能力。为增强稳定性与鲁棒性,LOTUS利用双模拟度量,该度量为LTL表示提供理论保障,包括行为等价性、最优保真度与轨迹鲁棒性。实验结果显示,LOTUS在学习效率、泛化能力与表示质量上优于多数现有方法:在单任务场景中,LOTUS使收敛速度提升超过20%;在未见过的操纵任务中,成功率提高15%-45%;在子目标深度或合取项增加的复杂多任务环境中,泛化性能提升超过25%。相关代码、视频与附录可在该网址获取。

英文摘要

Task guided agents demonstrate strong performance in a wide range of complex tasks. However, most existing task representation algorithms are tailored to specific contexts and struggle to generalize across diverse scenarios. Moreover, they typically depend on gradient signals from reinforcement learning controllers to update their weights, which can degrade both representation quality and learning efficiency. To overcome these limitations, we propose LOTUS, a temporal logic inspired universal task representation framework that can be seamlessly integrated into any RL algorithm to enhance agent performance across diverse task settings. Specifically, we design a novel task representation architecture capable of modeling relationships and extracting task semantics from LTL formulas. We further introduce a more effective update mechanism that treats the LTL encoder as a policy, thereby improving representation capacity. To enhance stability and robustness, LOTUS leverages the bisimulation metric, which provides theoretical guarantees for LTL representation, including behavioral equivalence, optimality fidelity, and trajectory robustness. Experimental results show that LOTUS outperforms most existing methods in learning efficiency, generalization capability, and representation quality. Specifically, LOTUS accelerates convergence over 20% in single-task scenarios, achieves a 15%-45% higher success rate in unseen manipulation tasks, and improves generalization performance over 25% in complex multi-task environments with increased sub-goal depth or conjunctions. The corresponding code, videos, and appendix are available at: https://lotus-website.github.io/.

CommentsAccepted by IEEE Transactions on Neural Networks and Learning Systems (Early Access). Project page: https://lotus-website.github.io/

DOI:10.1109/TNNLS.2026.3698967

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑