arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

前瞻性后见之明:通过预测-现实差距进行自校准强化学习

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

Jiaxin Zhang, Xiangyu Peng, Qinglin Chen, Yu Li, Hiroaki Hayashi, Chien-Sheng Wu

arXiv 2610.02740首次发表:更新:

发表机构

Salesforce AI Research(赛富时人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出前瞻性后见之明(PH)自校准训练原则,利用预测-现实差距加权梯度,在单轮与多轮任务中同时提升性能与校准,且适用于不同模型规模。

AI 中文摘要

针对长期智能体的强化学习依赖于纯粹的回顾性训练信号:仅在观察环境后果后分配信用,使得智能体在行动时的信念对梯度不可见。我们引入了前瞻性后见之明(Prospective Hindsight, PH),这是一种自校准训练原则,它通过智能体的前瞻性预测(反馈前)与回顾性评估(反馈后)之间的差距,为任何回顾性基础方法补充一个信号。这种每回合的意外度识别出智能体自模型最不准确的样本,并通过停止梯度的意外加权优势放大其梯度贡献。由于前瞻性预测器与策略共享参数,两者共同演化,逐步将焦点转移到智能体剩余的盲点上。我们将这一原则与特权信息差距联系起来,并表明最小化意外残差为智能体的误校准率提供了一条下降路径;因此,校准是优化的副产品,而非附加目标。在单轮可验证任务和多轮个人智能体任务(在GRPO、策略蒸馏及其组合下),PH同时提高了任务性能和校准,并在不同模型规模上保持一致增益。值得注意的是,主要的误校准模式在不同机制间结构性转变,单轮中表现为过度自信的失败,多轮中表现为信心不足的成功,但同一训练原则成功解决了这两种情况。

英文摘要

Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that augments any retrospective base method with a signal derived from the gap between the agent's prospective prediction (before feedback) and the retrospective evaluation (after feedback). This per-rollout surprise identifies samples where the agent's self-model is most inaccurate and amplifies their gradient contribution through a stop-gradient surprise-weighted advantage. Since the prospective predictor shares parameters with the policy, the two co-evolve, progressively shifting focus to the agent's remaining blind spots. We connect this principle to a privileged-information gap and show that minimizing the surprise residual provides a descent pathway on the agent's miscalibration rate; calibration thus emerges as a byproduct of optimization rather than from an added objective. On single-turn verifiable tasks and a multi-turn personal-agent task (under GRPO, on-policy distillation, and their combination), PH improves both task performance and calibration, with consistent gains across model scales. Notably, the dominant miscalibration mode shifts structurally between regimes, overconfident failures in single-turn, underconfident successes in multi-turn, yet the same training principle addresses both successfully.

CommentsNeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑