Latent Policy Steering through One-Step Flow Policies
通过一步流策略实现潜在策略引导
机构 * Department of Artificial Intelligence, Yonsei University(人工智能系,延世大学) ; Microsoft Research(微软研究院)
AI总结 本文提出Latent Policy Steering (LPS),通过一步流策略实现高保真的潜在策略改进,解决了离线强化学习中回报最大化与行为约束之间的权衡问题,实现了无需复杂超参数调整的鲁棒性能。
Comments Accepted to RSS 2026, Project Webpage : https://jellyho.github.io/LPS/