Retrospective In-Context Learning for Temporal Credit Assignment with Large Language Models
回顾上下文学习用于大语言模型中的时间信用分配
机构 * Carnegie Mellon University(卡内基梅隆大学) ; The University of Hong Kong(香港大学) ; Stanford University(斯坦福大学)
AI总结 本文提出利用大语言模型进行回顾上下文学习,以提高强化学习中时间信用分配的样本效率和泛化能力。
Comments Accepted to NeurIPS 2025