arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提示词活在弧线上:Fisher--Rao 坐标下的高斯课程用于高效滚动的GRPO

Prompts Live on an Arc: Gaussian Curricula in Fisher--Rao Coordinates for Rollout-Efficient GRPO

Mei Okonkwo, Pixel Nomand, Julian Berg, Elena Voss, Lena Park, Marcus Hale, Adrian Cho, Sofia Reyes

arXiv 2609.38018首次发表:更新:

发表机构

University of Wisconsin–Madison; University of Washington(威斯康星大学麦迪逊分校; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出ARCUS采样器,利用Fisher--Rao弧长坐标设计高斯课程,通过卡尔曼滤波选择信息组,提升GRPO效率,在数学推理基准上平均准确率提高2.8-2.9点,滚动减少48-57%。

AI 中文摘要

组相对策略优化(GRPO)仅从采样响应存在分歧的提示词中学习:完全正确或完全错误的组具有零奖励方差,不贡献梯度,但仍消耗其滚动。提示词选择方法通过将采样引导向中间通过率来减少这种浪费,但它们以原始通过率或logit坐标启发式地选择目标、其宽度和不确定性模型。我们表明,GRPO具有通过率的自然坐标:伯努利Fisher--Rao流形上的弧长$\psi=\arcsin\sqrt{p}$。在弧长中,预期的GRPO更新是均匀的,直到两个边界斜坡;零方差组的概率由宽度为$1/\sqrt{2G}$的两个高斯边界层界定;通过率证据具有恒定噪声;pass@$k$和pass$^k$目标的梯度是高斯函数,其中心和宽度由$k$以闭式形式给出。因此,GRPO的提示词课程是弧长中的高斯函数,选择其中心等同于选择目标。我们将这一观察转化为ARCUS,一个即插即用的采样器,在弧长中使用卡尔曼滤波器跟踪每个提示词,通过目标匹配的高斯核乘以信息组预测概率对提示词评分,仅保留信息组用于不变的GRPO更新,并将目标推向预测产量保持在最佳值小松弛范围内的最困难目标。在六个数学推理基准和三个骨干网络上,ARCUS将GRPO的平均准确率提高了2.8--2.9个百分点,将动态采样的平均准确率提高了1.1--1.2个百分点,同时生成的滚动比动态采样少48--57\\%。

英文摘要

Group relative policy optimization (GRPO) learns only from prompts whose sampled responses disagree: a group that is entirely correct or entirely incorrect has zero reward variance, contributes no gradient, and still consumes its rollouts. Prompt-selection methods reduce this waste by steering sampling toward intermediate pass rates, but they choose the target, its width, and the uncertainty model heuristically, in raw pass-rate or logit coordinates. We show that GRPO comes with a natural coordinate for pass rates: the arc length $ψ=\arcsin\sqrt{p}$ on the Bernoulli Fisher--Rao manifold. In arc length, the expected GRPO update is uniform up to two boundary ramps; the probability of a zero-variance group is bounded by two Gaussian boundary layers of width $1/\sqrt{2G}$; pass-rate evidence has constant noise; and the gradients of the pass@$k$ and pass$^k$ objectives are Gaussians whose center and width follow from $k$ in closed form. A prompt curriculum for GRPO is therefore a Gaussian in arc length, and choosing its center amounts to choosing the objective. We turn this observation into ARCUS, a drop-in sampler that tracks every prompt with a Kalman filter in arc length, scores prompts by an objective-matched Gaussian kernel times the predicted probability of an informative group, keeps only informative groups for the unchanged GRPO update, and paces the target toward the hardest objective whose predicted yield stays within a small slack of the best. Across six mathematical reasoning benchmarks and three backbones, ARCUS improves the average accuracy of GRPO by 2.8--2.9 points and that of dynamic sampling by 1.1--1.2 points, while generating 48--57\% fewer rollouts than dynamic sampling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑