arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24380cs.LGcs.AI

信息时间近端策略优化

Information-Time Proximal Policy Optimization

Yongcheng Zeng, Xinyu Cui, Yan Song, Guoqing Liu, Hongsheng Xin, Kaike Zhang, Cheng Deng, Kun Zhan, Jian Ying, Jian Zhao, Haifeng Zhang, Jun Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型推理中时间推进与信息流不均的问题,提出基于信息密度重参数化时间并自适应裁剪的 InfoPPO,在多个数学推理基准上优于基线。

中文摘要 AI 辅助

RLVR 已大幅提升了大语言模型的推理能力。然而,现有方法通常通过逐 token 生成来参数化马尔可夫决策过程中的时间推进,尽管自回归轨迹上的信息流高度不均匀。本文提出 InfoPPO,它使用信息密度而非原始 token 数量重新参数化时间推进。这种重新参数化在时间信用传播和策略更新两方面引入了共同的状态依赖结构。InfoPPO 恢复了长程推理中非平凡折扣的有效性,在保留有效视界收缩的同时,避免了长 token 序列上终端监督的过度衰减。此外,信息时间策略改进分析自然导出一个状态依赖的更新约束,我们通过自适应裁剪来实现它。通过将每个 token 位置的裁剪阈值调整为其对应状态的信息密度,该机制在保持近端控制的同时实现了更有针对性的策略更新。理论上,我们将性能差异和策略改进分析扩展到信息时间 MDP,在策略变化受信息密度调控时推导出策略改进下界。我们进一步通过将状态级信息密度与局部策略移动相关联,将一般信息时间分析与实际的大语言模型策略优化联系起来,同时为自适应更新机制提供理论依据。在 Qwen3 模型上的实验表明,在五个具有挑战性的竞赛级数学推理基准上,相对于强基线取得了持续提升。InfoPPO 还在非平凡折扣设置下保持稳定的准确率和响应长度,而 token 时间 PPO 在这些设置下性能恶化。

英文摘要

RLVR has substantially improved the reasoning capabilities of LLMs. However, existing methods typically parameterize temporal progression in the Markov Decision Process by token-by-token generation, despite the highly non-uniform information flow along autoregressive trajectories. In this paper, we propose InfoPPO, which reparameterizes temporal progression using information density rather than raw token count. This reparameterization induces a common state-dependent structure for both temporal credit propagation and policy updates. InfoPPO restores the effectiveness of non-trivial discounting in long-horizon reasoning, retaining effective-horizon contraction while avoiding excessive attenuation of terminal supervision over long token sequences. Moreover, the information-time policy-improvement analysis naturally leads to a state-dependent update constraint, which we implement through adaptive clipping. By adapting the clipping threshold at each token position to the information density of its corresponding state, this mechanism enables more targeted policy updates while preserving proximal control. Theoretically, we extend performance-difference and policy-improvement analyses to the information-time MDP, deriving a policy-improvement lower bound when policy changes are regulated by information density. We further connect the general information-time analysis to practical LLM policy optimization by relating state-wise information density to local policy movement, while also providing theoretical grounding for the adaptive update mechanism. Experiments on Qwen3 models demonstrate consistent gains over competitive baselines across five challenging competition-style mathematical reasoning benchmarks. InfoPPO also maintains stable accuracy and response length across non-trivial discount settings under which token-time PPO deteriorates.

发表机构

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • University College London AI Lab(伦敦大学学院人工智能实验室)
  • Microsoft Research AI4Science(微软研究院AI4Science)
  • Li Auto(理想汽车)
  • The University of Edinburgh(爱丁堡大学)
  • Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑